Skip to content

Posture & maturity

PostureService answers “what does this data estate look like, and how mature is it?”. It serves the cockpit’s Maturity view and grounds the other agents.

RPCWhat it does
GetHealthScoreServes the latest fetch’s posture snapshot. With no snapshot, it runs the deterministic posture analysis.
BuildDomainGraphThe fast half of a session fetch. It derives data domains and their relationships deterministically, with no LLM call.
AnalyzeEstateCausalA do-calculus intervention over the connected estate (see below).
ListIncidentsData-maturity assessments (DataMaturityAssessment).
GetContextSemantic context items for an agent type.

Health score and grade

Each domain gets a DomainHealth. The structural gauges come from computeStructuralMaturity, which is deterministic Go over a domain’s structure, not a model. Each gauge is on a 0 to 1 scale, and the cockpit displays it as a percentage.

GaugeHow it’s computed
quality_stability_pctSampled completeness. A sparse sample stays neutral at 0.75.
lineage_pctA base of 0.55, plus 0.35 if a join key is present, capped at 1.
freshness_pctA base of 0.6, plus 0.28 if a temporal column is present.
risk_pii_pctThe share of columns detected as PII, times 2.5, capped at 0.9.
policy_pctmax(0.4, 0.85 − 0.4 × risk_pii)
cost_efficiency_pctA base of 0.7, adjusted by row count.

The health value is (quality + lineage + freshness + policy + cost) / 5 − 0.15 × risk_pii, and it maps to a letter grade:

GradeHealth
A≥ 0.85
B≥ 0.72
C≥ 0.58
D≥ 0.45
Fbelow 0.45

The base values come from the instance’s risk-scoring configuration. issues lists the structural reasons, for example “no join key detected (lineage risk)”.

A failed scan returns an error, not a grade of F. With no connected data source, the local path answers FailedPrecondition: “connect a data source first”.

The empty_trip category is always empty. Nothing in the platform measures empty-trip or deadhead running, so the server clears it on every response, and the cockpit shows “Not measured — No backing record”.

Measure wording records where it came from. OptimizationSOP.provenance (MeasureProvenance) is one of:

  • MEASURED: read from a source.
  • DERIVED: computed in Go from measured data.
  • MODEL_GENERATED: written by the posture agent.

Only MODEL_GENERATED may be labelled as AI-generated.

Lineage

lineage_pct is a schema property, not observed lineage. It reads the same whether a source has been pulled a thousand times or never, because it depends only on whether a join key exists.

Observed lineage is a separate record, internal/lineage:

  • Where records come from. Every tabular connector is wrapped by one lineage decorator, so each extraction and each delivery is recorded. Camera ingests are recorded at their own seam.
  • How they’re stored. Records are append-only Avro in the object store, partitioned by UTC day.
  • What they contain. Structure only: column names, counts, durations, a literal-free query shape and a fingerprint, and a closed-set error class. Never result values, driver messages or credentials. A test with sentinel values enforces this.
  • How to read them. LineageService.GetSourceActivity returns per-connection activity over a window:
    • extractions, deliveries and rows.
    • The outcome mix: ok, empty and failed. empty means the source connected and returned nothing, which is neither a failure nor a delivery.
    • error_classes, destinations, and the last_columns names.
    • durable and durable_reason, which say whether the records survive a restart. Without a durable sink, the process keeps only a small in-memory ring.

Some reads aren’t in the log. A camera frame ingest is recorded with a row count of zero, because a camera sends frames, not rows. Live detection calls are not recorded. See Vision & CCTV.

The domain graph

BuildDomainGraph derives domains (structuralCommunities) and structural edges deterministically. It computes no correlations. Measured CORRELATED_WITH edges come from the fetch pipeline’s time-series EDA instead. See Fetch pipeline.

When connected sources are read, each edge’s relation names its evidence class. The classes are a closed set, listed from strongest to weakest:

  1. DECLARED_FOREIGN_KEY: the source’s own constraint.
  2. CATALOGUE_CONTAINMENT
  3. INFERRED_FROM_DECLARED_KEY
  4. INFERRED_SHARED_VOCABULARY
  5. INFERRED_FROM_VALUE_OVERLAP
  6. GUESSED_COLUMN_OVERLAP: a guess, and named as one.

Value sampling is on by default, because sampled values are real evidence about what a table is. metadata_only_connection_ids turns it off per connection. Sampled values reach only counts and shapes in the graph. They never appear in a log line, an error, a lineage record or a model prompt.

For the estate map, the cockpit now uses the shared estate component. FACE still serves include_connected_sources for older clients and for AnalyzeEstateCausal.

Estate causal analysis

AnalyzeEstateCausal runs an intervention over columns of the connected estate. A node is <qualified dataset>.<column>, carrying the source’s own aggregate over that column.

FACE separates structure from magnitude:

  • Structure. The estate may propose an edge only from a DECLARED_FOREIGN_KEY, and only where both columns have a measured value. No weaker evidence class proposes anything.
  • Magnitude. The coefficient is always the caller’s (EstateCausalCoefficient), because nothing in a catalogue or a sample measures how much one column moves another. Values that aren’t finite are refused.

How to read the response:

  • Read answered first. With no admitted edge, the RPC refuses and explains why in refusal. It never returns an effect of 0 with an OK status.
  • weakest_basis is the weakest basis among the edges the effect travelled along:
    • OPERATOR_ASSERTED: the caller named both endpoints.
    • DECLARED_KEY_REPURPOSED: the endpoints came from a declared key, which is weaker because a referential constraint is being read as a causal one.
  • basis_statement is written to be shown verbatim beside the number.
  • node_refusals, edge_refusals, nodes_added and columns_refused give the denominators. A graph of four columns out of nine hundred measurable ones is not a causal graph of the estate.

What this does not establish

  • A grade reflects structure, not how the data is used. A grade of A means the schema has keys, timestamps and few PII columns. It doesn’t mean the data is correct, current or governed.
  • lineage_pct is not lineage. Use GetSourceActivity for what was actually read and delivered.
  • A declared foreign key is a referential constraint, not a causal claim.
  • An estate causal answer reflects the coefficients you supplied. It rests on those coefficients, over the columns that had measurements.