Models & inference
Everything on the model plane is sovereign. It serves the platform’s own models on its own engine, and there is no hosted-model path. mistral.rs is the only engine (owner, 2026-09-27). The retired llama.cpp inference-engine image is not an option or a fallback.
Model cards (?tab=modelCards)
GET /api/models (models.go, session-gated) returns one row per serving tier, joined from three sources:
| Source | What it gives |
|---|---|
| card | modelcards.json, embedded in the console: provenance, licence, pin, intended use, known limits, evaluation evidence |
| live | the same tier projection /api/inference renders: model, quantization and context from the container spec; ready, restarts, last termination and digest from the pods; memory from metrics.k8s.io |
| bench | the core-inference-bench ConfigMap, attached only when the measurement can be attributed to that tier |
The health verdict is the join. A tier is healthy only if it is deployed, ready, not just OOM-killed, and running what its card names. A tier that serves happily from the wrong weights is not healthy. Live memory at or above 90% of the limit is degraded (modelMemoryWarnPct).
Usage is reported only as far as it exists:
- Token use comes from the app agents’ own
/api/agent-healthreports. It is cumulative since each agent process started, and is neither a time series nor per tenant. - Per-tenant compute is
ClientInstance.spec.computeUnits. - Admission, queue depth and token quotas are not exported by anything. They appear as
unmeasured[]entries.
Inference (?tab=inference)
The Inference page reads GET /api/models and the inference.router block of GET /api/agents. GET /api/inference (inference.go) remains the plane as deployed: engine and version (from the image reference; :latest is reported as :latest, not as a version), running digest, model and quantization (from --quantized-file/--model-id), context window, device, requests and limits, live memory and CPU, ready/desired, restarts and last termination.
- An empty context window means the engine default was taken. That is a finding, not a blank.
- Live usage is absent, not zero, when no metrics source is available. When metrics.k8s.io does not exist, the console reads the kubelet Summary API through the apiserver node proxy and labels the source
kubelet-summary. A 403 is unmeasured, never zero.
The model router (cmd/modelrouter)
One stdlib OpenAI-compatible front door that sends each request to a tier’s backend based on the request’s model field. It is also the one admission queue in front of the engines (admission.go).
| Env | Default | Meaning |
|---|---|---|
LISTEN_ADDR | :8080 | listen address |
BACKEND_GENERAL | required | the always-on general backend |
BACKEND_CODER, BACKEND_VISION, BACKEND_EMBEDDING, BACKEND_VOXTRAL | unset | optional tiers. Coder and vision fall back to general. Embedding and voxtral have no fallback, because a chat model returns garbage vectors or no audio support |
CONCURRENCY_<TIER> | 1 | slots per backend. 0 means no queue, direct dispatch |
QUEUE_DEPTH_<TIER> | 32 | queue length |
QUEUE_GROUP_<TIER> | the backend URL | tiers in one group share one queue |
QUEUE_MAX_WAIT | 15m | wait for a slot, then 503 |
RETRY_AFTER | 30s | the Retry-After value on a 503 |
INTERACTIVE_PER_PRINCIPAL | 2 | interactive requests one person may hold. Excess requests queue as background, never refused |
BACKGROUND_MAX_WAIT | 5m | background that has waited this long goes first, so agents are not starved |
UPSTREAM_HEADER_TIMEOUT | 20m | upstream response-header timeout |
CORE_ROUTER_IDENTITY_KEY | unset | key that verifies interactive claims |
How admission works:
- One queue per backend, not per tier. The general and coder tiers reach the same pods, so they share one slot count.
- A slot is held until the proxied body has been copied, because decoding is what occupies the engine. A client that disconnects cancels upstream and frees the slot.
- Interactive only with proof.
X-Runink-Priority: interactivecounts only when it arrives with anX-Core-Identityassertion (audiencemodel-router, signed withCORE_ROUTER_IDENTITY_KEY). The per-person cap counts the verified email. Every other priority claim is stripped. With no key, nobody can claim interactive, and everything is still routed. - A full queue answers
modelrouter: inference queue full; retry later. A wait pastQUEUE_MAX_WAITanswersmodelrouter: no inference slot within <wait>; retry later. Both are 503 withRetry-After. GET /statusreports each backend’s queue depth by class.- It orders only traffic that goes through it, and only with one replica (the queue is in-process).
The key is minted once by the core-operator’s provisioner into core-system/model-router-identity-key (provisioner/secrets.go). The console’s Deployment and the router’s onhost/12-model-router.yaml both reference it. The console’s model client (llmclient.go) signs every call made for a person waiting in the console. Such a call sends X-Runink-Priority: interactive, X-Runink-Principal (a digest, never the address) and X-Core-Identity. Background calls from agents and schedulers send none of the three. With no key configured, the assertion is omitted and the call is treated as background: slower, but never refused.
Two callers stay direct to the engine on purpose: opsdoctor, which must not depend on a hop it reports on, and eval bench, which measures the engine’s own TTFT and decode rate.
Budgets come from one rate
Every model-calling agent derives its max_tokens and timeouts from one configured decode rate, INFERENCE_DECODE_TOK_S (the inference library’s DecodeRateEnv). The default is 3.4 tok/s, the System76 estimate for Qwen3-14B, to be measured. The platform decodes one request at a time, so a 2000-token answer is roughly ten minutes of pure decode. Size budgets from the rate, never from wall-clock intuition, and never from a developer workstation.