Skip to content

Fetch pipeline

A fetch is one run of FACE’s analysis over the sources an operator has connected. It reads data, profiles and clusters it, runs the analysis agents, and returns a FetchAnalysisResult with decisions, action cards and evidence documents. All of this is served by FetchService.

RPCWhat it does
RunFetchRuns a fetch and returns one RunFetchResponse.
RunFetchStreamRuns the same fetch and streams RunFetchUpdate events as each phase finishes. The last event carries the full result.
ListTraces / GetTraceDetailsRead back the recorded execution history.
GenerateSOPDrafts a Standard Operating Procedure as markdown.
ListFetchStartersReturns suggested prompts for starting a fetch.
TalkBidirectional audio conversation with the fetch agent.

Schedules are a separate service, SchedulesService, and are covered below.

Where a fetch runs

The control plane never runs the pipeline itself. RunFetchStream on a control-plane process forwards the request to a dedicated runner and relays that runner’s stream back. If no runner can be dialled, the stream ends with a terminal final event that says the fetch did not run: “dedicated runner unavailable — fetch does not run on the control plane”. It never falls back to running the fetch locally.

Before any runner time is spent, the control plane checks the caller’s compute-unit budget. An exhausted budget ends the stream with a single event whose phase is denied. A successful fetch is then charged against that budget. Budgets are managed through MeteringService (see Operations).

The control plane resolves the connection records the fetch needs and sends them with the request in dispatched_connections. Credentials travel separately: each bundle is sealed to the runner’s ephemeral key for that one dispatch (dispatched_credentials, with dispatch_id). A client can’t set any of these fields, because the control plane overwrites them on the forward. How the sealing works is covered in Security.

Phases

    sequenceDiagram
    participant C as Cockpit
    participant CP as Control plane
    participant R as Runner
    participant M as Model plane
    C->>CP: RunFetchStream(RunFetchRequest)
    CP->>CP: compute-unit budget check
    CP->>R: forward (records + sealed credentials)
    R-->>C: phase=data_health (done=false)
    R->>R: read sources, profile, cluster, EDA
    par Phase 1
        R->>M: posture agent (TOON)
    and
        R->>M: rules reconciliation (TOON)
    end
    R-->>C: phase=data_health (done=true)
    R->>M: Phase 2 hypothesis swarm
    R-->>C: phase=scenarios
    R->>M: Phase 3 twins decision
    R->>R: judging ladder on proposed actions
    R-->>C: phase=recommendations
    R-->>C: phase=final + RunFetchResponse
  

RunFetchUpdate.phase takes these values on the stream: data_health, scenarios, recommendations and final, plus denied for a budget refusal. done=false means a phase is still running. The terminal event always has phase="final", done=true, and a result.

The runner’s pipeline (runContextEngineeringPipeline) has three phases:

  1. Phase 1, data alignment. The posture agent and the rules-reconciliation agent run in parallel over the fetched context. Before they run, the pipeline clusters the samples (HDBSCAN through AnalysisService.ClusterRefinement) and runs a per-domain time-series EDA (runDomainTimeSeriesEDA, backed by ml/prophet). When two data domains are correlated with |r| at or above RunFetchRequest.synapse_threshold (default 0.5), the EDA records a measured CORRELATED_WITH edge between them.
  2. Phase 2, hypothesis swarm. A bounded-parallel swarm runs one agent per scenario the estate proposes. The swarm has a cap, and its census travels in FetchAnalysisResult. See Simulation & twins.
  3. Phase 3, twins decision. The twins orchestrator turns the grounded findings and simulations into decisions and action cards. Each proposed action goes through the judging ladder before it’s accepted. See Agents & oversight.

Session fetches are lighter. When session_eda is set, the fetch runs Phase 1 only: data domains, maturity, rules and the knowledge graph. It never runs the hypothesis or twins agents. The cockpit uses this for its automatic session fetch. A manual, scheduled or API fetch runs the full pipeline.

Unary RunFetch returns early. When Phase 1 completes, RunFetch returns the core result, and Phases 2 and 3 keep running in the background. RunFetchStream waits until the recommendations are ready.

What the result carries

FetchAnalysisResult includes the following fields, among others:

  • decisions and action_cards: the structured twin output. See Action cards & money.
  • graph_nodes and graph_edges: the knowledge graph for the fetched domains.
  • hypothesis_scenarios, hypothesis_impacts, hypothesis_traces, hypothesis_forecasts and hypothesis_anomalies: the swarm’s union of answers. It is a union, not a consensus. Nothing votes on or reconciles the agents.
  • hypothesis_scenarios_inferred, hypothesis_agents_dispatched and hypothesis_agents_succeeded: the swarm census. These counts are only meaningful when hypothesis_swarm_ran is true.
  • source_notices: sentences about sources that could not be read, or about where a source’s SQL actually ran. An empty list does not mean every source was fine.
  • action_judgements: the judging ladder’s verdict on every proposed action, whether accepted or rejected. A rejected action appears nowhere else.
  • provenance: how the answer was produced, fresh or exact-cache, set by the server.
  • metrics (EvaluationMetrics): these scores are not filled in by the fetch. The response sets unmeasured_reason to “no evaluation harness ran for this fetch — these scores are absent, not zero”. A client must not render the scores as percentages when that field is set.
  • sql_generated: SQL text drafted by the model from the fetched schemas. It’s informational output, not a record of what was executed.

Repeat fetches can be served from cache. A recent identical fetch (same connection_id and command) can be replayed from cache. The cache entry must be younger than FETCH_CACHE_TTL (default 10 minutes) and at least FETCH_CACHE_MIN_CONFIDENCE (default 0.7) confident. A cached answer is marked in provenance. A run whose analysis phases all failed is never cached.

Traces

ListTraces and GetTraceDetails read the session records that each fetch saves to the durable record store. A Trace from these RPCs carries the transaction id, the command, the analysis text and the drafted SQL. If the process has no session store, ListTraces answers FailedPrecondition instead of returning an empty list, because “never read” and “nothing recorded” are different answers.

Observed data lineage, meaning which source was read when and where the rows went, is a different record. It’s served by LineageService.GetSourceActivity. See Posture & maturity.

SOP generation

GenerateSOP takes a topic, a process_name and optional steps. The prompt is checked in two places before it reaches the model:

  • ai.SanitizePrompt rejects the prompt with InvalidArgument, and the refusal is audit-logged.
  • The FetchService OpenBias guardrails.

The response carries markdown in sop_content and an sop_id. If the model call fails after the prompt is accepted, the content is the literal SOP Generation Failed.

Fetch starters

ListFetchStarters returns up to four FetchStarter cards, each with an icon, a title, a subtitle and prompt_text. The cards are sampled from the instance’s configured fetch_starters, or from a built-in pool when none are configured. They are prompt suggestions and carry no data.

Schedules

SchedulesService provides ListSchedules, CreateSchedule, UpdateSchedule, SetScheduleActive and DeleteSchedule. A Schedule has a free-text frequency, which maps to an interval: hourly, every 6 hours, daily, weekly or monthly. Unrecognised text maps to hourly. The executor runs on the control plane, not on runners.

  • Execution is opt-in. A due schedule runs a real fetch only when SCHEDULE_EXECUTE_FETCH=true. Otherwise it is recorded as a skipped dry run, not as a success, and run_count does not increase.
  • Schedules run as a service account. A scheduled fetch runs as TRIGGER_TYPE_SCHEDULED under a short-lived service-account session (ROLE_SERVICE_ACCOUNT) minted for the scheduler.
  • Outcomes are recorded as they happened. last_run_success reflects what the fetch actually returned. A fetch declined by the pipeline, for example for budget, is recorded as a failure. A schedule whose previous run hasn’t finished skips that interval instead of stacking.

What this does not establish

  • A completed fetch is not a complete fetch. Read source_notices and the swarm census before treating the result as covering every selected source and scenario.
  • Progress events report pipeline stages, not answer quality. A final event with success=true means the pipeline finished. It doesn’t mean the model’s answer is correct.
  • EvaluationMetrics is not a quality score while unmeasured_reason is set.
  • A cached result is a replay of an earlier run. Its provenance.computed_at says when it was produced.