Ask
Ask is a grounded, citing question box over what CORE already holds, for the Intelligence and DataEx categories. It implements the owner’s 2026-09-26 decision that CORE’s Intelligence and DataEx reach FACE’s level of reasoning and grounding.
- Contract:
AskServiceingrpc/operators/core/api/proto/runink/core/ask/v1/ask.proto - Server code:
internal/console/ask.go,ask_tools.go,ask_facts.go,ask_cite.go,ask_guard.go,llmclient.go
AskService on its in-process gRPC-web server, and the Dart client code is generated (flutter/lib/core/grpc/runink/core/ask/v1/). No console page calls it yet: nothing under flutter/lib/features/ imports it.The two RPCs
| RPC | Class | What it does |
|---|---|---|
Describe | read | Lists, per scope, the tools the model may call and the limits a run is held to, plus the active guardrail profile |
Ask | streaming, gated by consoleStreamInterceptor | Runs one grounded question. It streams the loop’s steps as they happen and ends with exactly one Answer, or with a status |
Ask is the console’s only streaming RPC. It requires a session on a console with sign-in, because the run acts for a person: the tool gate, the interactive-priority principal and the audit actor all need one. An open console refuses it with this console has no sign-in configured, so an Ask would act for nobody — … Set GOOGLE_CLIENT_ID or CONSOLE_ADMIN_PASSWORD to use Ask.
scope must be SCOPE_INTELLIGENCE or SCOPE_DATAEX. A question is at most 2,000 bytes.
What runs
One Ask is one inference/agent.Run. That is FACE’s ReAct loop, extracted into the shared inference library and reused here rather than rewritten. It runs on the console’s one model client (callLLMMessages over INFERENCE_REMOTE_URL, which is the model router). The model is sovereign only, never a hosted model.
If INFERENCE_REMOTE_URL is unset, the stream ends with UNAVAILABLE: inference not configured (INFERENCE_REMOTE_URL unset) — no model was called.
The budget is derived, not declared
The budget follows the platform’s one configured decode rate, INFERENCE_DECODE_TOK_S. When that is unset, the System76 default of 3.4 tok/s applies.
| Limit | Value |
|---|---|
| Steps per run | 4 (up to three lookups and the answer) |
| Completion tokens per run | 1,800 (about 9 minutes of pure decode at 3.4 tok/s) |
| Tokens per turn | at most 640, never fewer than 320 |
| One tool result shown to the model | 2,400 bytes |
The deadline is the sum of three terms, with no padding:
- the decode time for the completion tokens;
- one prefill per step;
- the wait for one turn already decoding.
At the default rate this comes to minutes, which is why Ask streams its steps.
Tools are read-only, in-process, and never dial
The model may call only read-only methods that the console already serves, called in process:
- the Atlas gRPC reads, on the same service structs;
- the JSON reports, through their own handlers with the caller’s session cookie.
There is no second data path. No tool dials a tenant source. Resolve’s four dialing actions are not tools. Every tool output enters the prompt fenced by prompt.FenceUntrusted.
| Scope | Tools (ask_tools.go) |
|---|---|
| Intelligence | resolve_sources, resolve_source_description, resolve_estate, resolve_reconciliation, capex_summary, capex_findings, capex_trace, capex_rulebook, datagov_findings, governance_rules, governance_remediations, playbook_runs, data_lineage |
| DataEx | model_cards, agents, agent_config, guardrails, harness_findings, judgements, inference_status |
Bounded row-sample and SELECT tools are planned for the extension point askRowToolsExtension, which is nil in this build on purpose. It is never a stub that answers “no rows”.
Grounding is enforced
- Every figure in the answer must be a value some tool result of this run showed.
ask_facts.gobuilds the fact table from exactly the lines that were rendered to the model. After the model finishes, the server checks every number. - A figure that survives carries a
Citation:{kind, record_id, field, value as shown, tool}. - Anything else, including any arithmetic the model did, is cut from the text and listed in
removed_figures. - Each lookup reports an outcome.
COULD_NOT_LOOK(a store error, provenance UNAVAILABLE) is neverFOUND_NOTHING, on the wire, to the model or in the caveats.REFUSEDis its own outcome. - An answer is
ANSWEREDorSTOPPED.STOPPEDmeans a limit was hit (steps, budget, no progress, deadline) and comes with astop_reason. - No model-authored score exists on this surface: no confidence, quality score or PII score, and the contract has no field for one.
Guardrails
The console builds ONE security/guardrails Guard at start, with profile console-ask. If it cannot be built, the console refuses to start.
- What the person typed is screened at input, before any model call.
- The model’s text is screened at output, before it is shown or acted on. Model “thoughts” are screened too, and a blocked thought is replaced by a notice.
- The assembled prompt is not screened. The fence covers fetched content instead.
- A block ends the stream with
PERMISSION_DENIEDand the guardrail’s own wording. Nothing of the answer is sent. The same Guard screens/chat, where a block is an HTTP 403.
Audit records:
- Guardrail refusals are recorded under
guardrails.<action>(for exampleguardrails.prompt_rejectedandguardrails.output_rejected), with only the hashed actor and the rule ids. - Each Ask is recorded as
ask.<scope>under resourceconsole-ask, with the question’s hash, the scope, the tools used, the outcome and counts. It never records the question or answer text.
The Model guardrails rail on /api/security and modelGuardrails on /api/guardrails (DataEx › Trust › Guardrails) are read from the Guard itself, never from env.
The interactive lane
A call made for a person sends three headers:
X-Runink-Priority: interactive;X-Runink-Principal, a digest and never the address;X-Core-Identity, an assertion for audiencemodel-router, signed withCORE_ROUTER_IDENTITY_KEY(Secretcore-system/model-router-identity-key, minted by the provisioner and mirrored into the inference namespace).
The router honours interactive priority only with a verified assertion. With no key configured, the assertion is omitted and the router treats the call as background: slower, but never refused. Background calls (agents, schedulers) send none of the three headers. See Models & inference.