Statistics & ML
FACE’s classical analysis comes from the Runink ml library. It’s pure Go, involves
no neural network or external model server, and is standard-library only except for
ml/linear and ml/segment. FACE exposes this analysis through AnalysisService, and
the fetch pipeline uses the same code internally.
AnalysisService
AnalyzeTimeSeries
Forecasts a series forecast_horizon steps ahead. The default horizon is 3.
- Minimum input. Fewer than 5 points returns
recommended_model="Insufficient Data (Need > 5 points)"and no forecast. - Primary path:
prophet.AnalyzeSeries. This is a full series EDA. It runs descriptive statistics, then an ADF stationarity test, then ACF/PACF autocorrelation, then a Prophet decomposition (piecewise-linear trend, Fourier seasonality and changepoints). Finally it fits ARIMA, and a rolling backtest picks between ARIMA and Prophet. The response carries:recommended_model: the chosen model, marked[backtest-selected], for exampleARIMA(p,d,q) [backtest-selected].stationarity_test: the ADF result, with the statistic and the 5% critical value.anomalies: the EDA’s findings, plus any Prophet residual anomalies.
- Fallback. If the EDA can’t fit a model, an ordinary-least-squares line is
extrapolated. The response says
Linear Regression (Fallback)and states that the EDA was unavailable.
AnalyzeFeature
Tests autocorrelation at lags 1 to 5 of target. Any lag with |ACF| > 0.5 is
reported in significant_lags, and a target_lag_k feature is proposed for it.
Fewer than 10 points returns Insufficient data, and a constant series returns
Constant target.
ClusterRefinement
Clusters embedded items into domains:
- Validate the vectors. The first non-empty vector sets the dimension. Vectors
that are missing or have a different dimension are kept in place as noise
(
-1) rather than dropped, so labels stay aligned with the caller’s items. - Cluster with HDBSCAN.
knowledge.SelectHDBSCANruns with a minimum cluster size of 3 and Euclidean distance. It sweeps the minimum cluster size and keeps the candidate with the best silhouette coefficient. Candidates with fewer than two clusters, or with more than half the points in noise, are skipped. - Rerank. A heuristic reranker in
ml/xgboostapplies per-domain feature weights, using BM25 relevance of each item against the request’s metadata as an extra feature. Despite the package name, this step doesn’t use a trained model.
The response returns clusters (a label per item), noise (the indices labelled
noise) and health_score. The health score is the silhouette, or the reranker’s
cluster health when reranking succeeded.
AnalyzeCausal
A do-calculus intervention over a graph the caller states in full:
nodesmaps each node to its observed value.edges(CausalEdge) carry a structuraleffect_size.- An intervention sets
intervention_nodetointervention_value, and the change propagates along edges by multiplying through each effect size. - The response returns
effectontarget_nodeand the target’snew_target_value. - An edge naming an unknown node, or an edge set with a cycle, is refused with
InvalidArgument. The call never runs over a partly built graph.
An edge at 0 declares that there is no relationship. A missing edge means the
relationship is unknown. To run an intervention over the connected estate instead
of a hand-written graph, use PostureService.AnalyzeEstateCausal. See
Posture & maturity.
InferBayesian
Posterior inference over a Bayesian network (ml/knowledge). The request carries
evidence and a target node, but it has no field for defining variables or conditional
probability tables. So on an instance with no defined network, the call answers
FailedPrecondition and explains why, instead of returning an empty result with an
OK status.
AnalyzeDescriptiveStats
Returns the mean, median, variance, standard deviation, skewness and kurtosis of
data, computed by knowledge.DescriptiveStats.
ClassifyESGRelevance
Returns relevance_score, a keyword-lexicon score from 0 to 1. Each text is scored
by how often terms such as carbon, emission, net-zero and governance appear,
weighted per term and capped at 1.0. It’s a relevance heuristic, not a trained
classifier. The same texts are also passed to FACE’s digestion pipeline (clustering,
EDA, forecast and causal passes). A failure there is logged and doesn’t change the
score.
What ml provides underneath
| Package | Used by FACE for |
|---|---|
ml/knowledge | HDBSCAN and silhouette selection, k-means, BM25, causal do-calculus, Bayesian networks, descriptive statistics, community detection and domain inference. |
ml/prophet | Prophet decomposition, ARIMA, ADF, ACF/PACF and the unified series EDA. |
ml/xgboost | Two separate things. The heuristic cluster reranker above, and a real gradient-boosted tree implementation with probability calibration, which ml/decide uses. |
ml/decide | The pinned tree-model decision engine behind the judging ladder’s fast path. See Agents & oversight. |
ml/segment | Region proposals for vision. It pools a frame into grid features, which are then clustered. It proposes where things are, never what they are. |
ml/linear | Linear regression on a CPU-only backend. FACE wires the library’s audit hook. |
ml holds no state. Everything is computed and returned, and persistence is the
calling app’s concern.
What this does not establish
- A forecast is a model’s extrapolation of the series you supplied. It knows nothing outside that series. The backtest picks the better of two models on the series’ own history, and that is not a guarantee about the future.
- A causal effect is the consequence of the graph you stated. The server computes it; it doesn’t discover it. Wrong edges give wrong effects, with no error.
- Correlation edges are not causal edges. Measured correlations between data
domains in a fetch are recorded as
CORRELATED_WITHand nothing more. - A cluster is a statistical grouping. A high silhouette means the points separate well. It doesn’t mean the domain names attached to them are right.
- The ESG score counts words. It doesn’t assess performance against any ESG standard.