Skip to content

Statistics & ML

FACE’s classical analysis comes from the Runink ml library. It’s pure Go, involves no neural network or external model server, and is standard-library only except for ml/linear and ml/segment. FACE exposes this analysis through AnalysisService, and the fetch pipeline uses the same code internally.

AnalysisService

AnalyzeTimeSeries

Forecasts a series forecast_horizon steps ahead. The default horizon is 3.

  • Minimum input. Fewer than 5 points returns recommended_model="Insufficient Data (Need > 5 points)" and no forecast.
  • Primary path: prophet.AnalyzeSeries. This is a full series EDA. It runs descriptive statistics, then an ADF stationarity test, then ACF/PACF autocorrelation, then a Prophet decomposition (piecewise-linear trend, Fourier seasonality and changepoints). Finally it fits ARIMA, and a rolling backtest picks between ARIMA and Prophet. The response carries:
    • recommended_model: the chosen model, marked [backtest-selected], for example ARIMA(p,d,q) [backtest-selected].
    • stationarity_test: the ADF result, with the statistic and the 5% critical value.
    • anomalies: the EDA’s findings, plus any Prophet residual anomalies.
  • Fallback. If the EDA can’t fit a model, an ordinary-least-squares line is extrapolated. The response says Linear Regression (Fallback) and states that the EDA was unavailable.

AnalyzeFeature

Tests autocorrelation at lags 1 to 5 of target. Any lag with |ACF| > 0.5 is reported in significant_lags, and a target_lag_k feature is proposed for it. Fewer than 10 points returns Insufficient data, and a constant series returns Constant target.

ClusterRefinement

Clusters embedded items into domains:

  1. Validate the vectors. The first non-empty vector sets the dimension. Vectors that are missing or have a different dimension are kept in place as noise (-1) rather than dropped, so labels stay aligned with the caller’s items.
  2. Cluster with HDBSCAN. knowledge.SelectHDBSCAN runs with a minimum cluster size of 3 and Euclidean distance. It sweeps the minimum cluster size and keeps the candidate with the best silhouette coefficient. Candidates with fewer than two clusters, or with more than half the points in noise, are skipped.
  3. Rerank. A heuristic reranker in ml/xgboost applies per-domain feature weights, using BM25 relevance of each item against the request’s metadata as an extra feature. Despite the package name, this step doesn’t use a trained model.

The response returns clusters (a label per item), noise (the indices labelled noise) and health_score. The health score is the silhouette, or the reranker’s cluster health when reranking succeeded.

AnalyzeCausal

A do-calculus intervention over a graph the caller states in full:

  • nodes maps each node to its observed value.
  • edges (CausalEdge) carry a structural effect_size.
  • An intervention sets intervention_node to intervention_value, and the change propagates along edges by multiplying through each effect size.
  • The response returns effect on target_node and the target’s new_target_value.
  • An edge naming an unknown node, or an edge set with a cycle, is refused with InvalidArgument. The call never runs over a partly built graph.

An edge at 0 declares that there is no relationship. A missing edge means the relationship is unknown. To run an intervention over the connected estate instead of a hand-written graph, use PostureService.AnalyzeEstateCausal. See Posture & maturity.

InferBayesian

Posterior inference over a Bayesian network (ml/knowledge). The request carries evidence and a target node, but it has no field for defining variables or conditional probability tables. So on an instance with no defined network, the call answers FailedPrecondition and explains why, instead of returning an empty result with an OK status.

AnalyzeDescriptiveStats

Returns the mean, median, variance, standard deviation, skewness and kurtosis of data, computed by knowledge.DescriptiveStats.

ClassifyESGRelevance

Returns relevance_score, a keyword-lexicon score from 0 to 1. Each text is scored by how often terms such as carbon, emission, net-zero and governance appear, weighted per term and capped at 1.0. It’s a relevance heuristic, not a trained classifier. The same texts are also passed to FACE’s digestion pipeline (clustering, EDA, forecast and causal passes). A failure there is logged and doesn’t change the score.

What ml provides underneath

PackageUsed by FACE for
ml/knowledgeHDBSCAN and silhouette selection, k-means, BM25, causal do-calculus, Bayesian networks, descriptive statistics, community detection and domain inference.
ml/prophetProphet decomposition, ARIMA, ADF, ACF/PACF and the unified series EDA.
ml/xgboostTwo separate things. The heuristic cluster reranker above, and a real gradient-boosted tree implementation with probability calibration, which ml/decide uses.
ml/decideThe pinned tree-model decision engine behind the judging ladder’s fast path. See Agents & oversight.
ml/segmentRegion proposals for vision. It pools a frame into grid features, which are then clustered. It proposes where things are, never what they are.
ml/linearLinear regression on a CPU-only backend. FACE wires the library’s audit hook.

ml holds no state. Everything is computed and returned, and persistence is the calling app’s concern.

What this does not establish

  • A forecast is a model’s extrapolation of the series you supplied. It knows nothing outside that series. The backtest picks the better of two models on the series’ own history, and that is not a guarantee about the future.
  • A causal effect is the consequence of the graph you stated. The server computes it; it doesn’t discover it. Wrong edges give wrong effects, with no error.
  • Correlation edges are not causal edges. Measured correlations between data domains in a fetch are recorded as CORRELATED_WITH and nothing more.
  • A cluster is a statistical grouping. A high silhouette means the points separate well. It doesn’t mean the domain names attached to them are right.
  • The ESG score counts words. It doesn’t assess performance against any ESG standard.