Skip to content
Models & inference

Models & inference

Everything on the model plane is sovereign. It serves the platform’s own models on its own engine, and there is no hosted-model path. mistral.rs is the only engine (owner, 2026-09-27). The retired llama.cpp inference-engine image is not an option or a fallback.

Model cards (?tab=modelCards)

GET /api/models (models.go, session-gated) returns one row per serving tier, joined from three sources:

SourceWhat it gives
cardmodelcards.json, embedded in the console: provenance, licence, pin, intended use, known limits, evaluation evidence
livethe same tier projection /api/inference renders: model, quantization and context from the container spec; ready, restarts, last termination and digest from the pods; memory from metrics.k8s.io
benchthe core-inference-bench ConfigMap, attached only when the measurement can be attributed to that tier

The health verdict is the join. A tier is healthy only if it is deployed, ready, not just OOM-killed, and running what its card names. A tier that serves happily from the wrong weights is not healthy. Live memory at or above 90% of the limit is degraded (modelMemoryWarnPct).

Usage is reported only as far as it exists:

  • Token use comes from the app agents’ own /api/agent-health reports. It is cumulative since each agent process started, and is neither a time series nor per tenant.
  • Per-tenant compute is ClientInstance.spec.computeUnits.
  • Admission, queue depth and token quotas are not exported by anything. They appear as unmeasured[] entries.

Inference (?tab=inference)

The Inference page reads GET /api/models and the inference.router block of GET /api/agents. GET /api/inference (inference.go) remains the plane as deployed: engine and version (from the image reference; :latest is reported as :latest, not as a version), running digest, model and quantization (from --quantized-file/--model-id), context window, device, requests and limits, live memory and CPU, ready/desired, restarts and last termination.

  • An empty context window means the engine default was taken. That is a finding, not a blank.
  • Live usage is absent, not zero, when no metrics source is available. When metrics.k8s.io does not exist, the console reads the kubelet Summary API through the apiserver node proxy and labels the source kubelet-summary. A 403 is unmeasured, never zero.

The model router (cmd/modelrouter)

One stdlib OpenAI-compatible front door that sends each request to a tier’s backend based on the request’s model field. It is also the one admission queue in front of the engines (admission.go).

EnvDefaultMeaning
LISTEN_ADDR:8080listen address
BACKEND_GENERALrequiredthe always-on general backend
BACKEND_CODER, BACKEND_VISION, BACKEND_EMBEDDING, BACKEND_VOXTRALunsetoptional tiers. Coder and vision fall back to general. Embedding and voxtral have no fallback, because a chat model returns garbage vectors or no audio support
CONCURRENCY_<TIER>1slots per backend. 0 means no queue, direct dispatch
QUEUE_DEPTH_<TIER>32queue length
QUEUE_GROUP_<TIER>the backend URLtiers in one group share one queue
QUEUE_MAX_WAIT15mwait for a slot, then 503
RETRY_AFTER30sthe Retry-After value on a 503
INTERACTIVE_PER_PRINCIPAL2interactive requests one person may hold. Excess requests queue as background, never refused
BACKGROUND_MAX_WAIT5mbackground that has waited this long goes first, so agents are not starved
UPSTREAM_HEADER_TIMEOUT20mupstream response-header timeout
CORE_ROUTER_IDENTITY_KEYunsetkey that verifies interactive claims

How admission works:

  • One queue per backend, not per tier. The general and coder tiers reach the same pods, so they share one slot count.
  • A slot is held until the proxied body has been copied, because decoding is what occupies the engine. A client that disconnects cancels upstream and frees the slot.
  • Interactive only with proof. X-Runink-Priority: interactive counts only when it arrives with an X-Core-Identity assertion (audience model-router, signed with CORE_ROUTER_IDENTITY_KEY). The per-person cap counts the verified email. Every other priority claim is stripped. With no key, nobody can claim interactive, and everything is still routed.
  • A full queue answers modelrouter: inference queue full; retry later. A wait past QUEUE_MAX_WAIT answers modelrouter: no inference slot within <wait>; retry later. Both are 503 with Retry-After.
  • GET /status reports each backend’s queue depth by class.
  • It orders only traffic that goes through it, and only with one replica (the queue is in-process).

The key is minted once by the core-operator’s provisioner into core-system/model-router-identity-key (provisioner/secrets.go). The console’s Deployment and the router’s onhost/12-model-router.yaml both reference it. The console’s model client (llmclient.go) signs every call made for a person waiting in the console. Such a call sends X-Runink-Priority: interactive, X-Runink-Principal (a digest, never the address) and X-Core-Identity. Background calls from agents and schedulers send none of the three. With no key configured, the assertion is omitted and the call is treated as background: slower, but never refused.

Two callers stay direct to the engine on purpose: opsdoctor, which must not depend on a hop it reports on, and eval bench, which measures the engine’s own TTFT and decode rate.

Budgets come from one rate

Every model-calling agent derives its max_tokens and timeouts from one configured decode rate, INFERENCE_DECODE_TOK_S (the inference library’s DecodeRateEnv). The default is 3.4 tok/s, the System76 estimate for Qwen3-14B, to be measured. The platform decodes one request at a time, so a 2000-token answer is roughly ten minutes of pure decode. Size budgets from the rate, never from wall-clock intuition, and never from a developer workstation.