Skip to content

Ask

Ask is a grounded, citing question box over what CORE already holds, for the Intelligence and DataEx categories. It implements the owner’s 2026-09-26 decision that CORE’s Intelligence and DataEx reach FACE’s level of reasoning and grounding.

  • Contract: AskService in grpc/operators/core/api/proto/runink/core/ask/v1/ask.proto
  • Server code: internal/console/ask.go, ask_tools.go, ask_facts.go, ask_cite.go, ask_guard.go, llmclient.go
Status: the console serves AskService on its in-process gRPC-web server, and the Dart client code is generated (flutter/lib/core/grpc/runink/core/ask/v1/). No console page calls it yet: nothing under flutter/lib/features/ imports it.

The two RPCs

RPCClassWhat it does
DescribereadLists, per scope, the tools the model may call and the limits a run is held to, plus the active guardrail profile
Askstreaming, gated by consoleStreamInterceptorRuns one grounded question. It streams the loop’s steps as they happen and ends with exactly one Answer, or with a status

Ask is the console’s only streaming RPC. It requires a session on a console with sign-in, because the run acts for a person: the tool gate, the interactive-priority principal and the audit actor all need one. An open console refuses it with this console has no sign-in configured, so an Ask would act for nobody — … Set GOOGLE_CLIENT_ID or CONSOLE_ADMIN_PASSWORD to use Ask.

scope must be SCOPE_INTELLIGENCE or SCOPE_DATAEX. A question is at most 2,000 bytes.

What runs

One Ask is one inference/agent.Run. That is FACE’s ReAct loop, extracted into the shared inference library and reused here rather than rewritten. It runs on the console’s one model client (callLLMMessages over INFERENCE_REMOTE_URL, which is the model router). The model is sovereign only, never a hosted model.

If INFERENCE_REMOTE_URL is unset, the stream ends with UNAVAILABLE: inference not configured (INFERENCE_REMOTE_URL unset) — no model was called.

The budget is derived, not declared

The budget follows the platform’s one configured decode rate, INFERENCE_DECODE_TOK_S. When that is unset, the System76 default of 3.4 tok/s applies.

LimitValue
Steps per run4 (up to three lookups and the answer)
Completion tokens per run1,800 (about 9 minutes of pure decode at 3.4 tok/s)
Tokens per turnat most 640, never fewer than 320
One tool result shown to the model2,400 bytes

The deadline is the sum of three terms, with no padding:

  1. the decode time for the completion tokens;
  2. one prefill per step;
  3. the wait for one turn already decoding.

At the default rate this comes to minutes, which is why Ask streams its steps.

Tools are read-only, in-process, and never dial

The model may call only read-only methods that the console already serves, called in process:

  • the Atlas gRPC reads, on the same service structs;
  • the JSON reports, through their own handlers with the caller’s session cookie.

There is no second data path. No tool dials a tenant source. Resolve’s four dialing actions are not tools. Every tool output enters the prompt fenced by prompt.FenceUntrusted.

ScopeTools (ask_tools.go)
Intelligenceresolve_sources, resolve_source_description, resolve_estate, resolve_reconciliation, capex_summary, capex_findings, capex_trace, capex_rulebook, datagov_findings, governance_rules, governance_remediations, playbook_runs, data_lineage
DataExmodel_cards, agents, agent_config, guardrails, harness_findings, judgements, inference_status

Bounded row-sample and SELECT tools are planned for the extension point askRowToolsExtension, which is nil in this build on purpose. It is never a stub that answers “no rows”.

Grounding is enforced

  • Every figure in the answer must be a value some tool result of this run showed. ask_facts.go builds the fact table from exactly the lines that were rendered to the model. After the model finishes, the server checks every number.
  • A figure that survives carries a Citation: {kind, record_id, field, value as shown, tool}.
  • Anything else, including any arithmetic the model did, is cut from the text and listed in removed_figures.
  • Each lookup reports an outcome. COULD_NOT_LOOK (a store error, provenance UNAVAILABLE) is never FOUND_NOTHING, on the wire, to the model or in the caveats. REFUSED is its own outcome.
  • An answer is ANSWERED or STOPPED. STOPPED means a limit was hit (steps, budget, no progress, deadline) and comes with a stop_reason.
  • No model-authored score exists on this surface: no confidence, quality score or PII score, and the contract has no field for one.

Guardrails

The console builds ONE security/guardrails Guard at start, with profile console-ask. If it cannot be built, the console refuses to start.

  • What the person typed is screened at input, before any model call.
  • The model’s text is screened at output, before it is shown or acted on. Model “thoughts” are screened too, and a blocked thought is replaced by a notice.
  • The assembled prompt is not screened. The fence covers fetched content instead.
  • A block ends the stream with PERMISSION_DENIED and the guardrail’s own wording. Nothing of the answer is sent. The same Guard screens /chat, where a block is an HTTP 403.

Audit records:

  • Guardrail refusals are recorded under guardrails.<action> (for example guardrails.prompt_rejected and guardrails.output_rejected), with only the hashed actor and the rule ids.
  • Each Ask is recorded as ask.<scope> under resource console-ask, with the question’s hash, the scope, the tools used, the outcome and counts. It never records the question or answer text.

The Model guardrails rail on /api/security and modelGuardrails on /api/guardrails (DataEx › Trust › Guardrails) are read from the Guard itself, never from env.

The interactive lane

A call made for a person sends three headers:

  • X-Runink-Priority: interactive;
  • X-Runink-Principal, a digest and never the address;
  • X-Core-Identity, an assertion for audience model-router, signed with CORE_ROUTER_IDENTITY_KEY (Secret core-system/model-router-identity-key, minted by the provisioner and mirrored into the inference namespace).

The router honours interactive priority only with a verified assertion. With no key configured, the assertion is omitted and the router treats the call as background: slower, but never refused. Background calls (agents, schedulers) send none of the three headers. See Models & inference.