Skip to content
Sovereign inference

Sovereign inference

Every AI feature in PULSE runs on the Runink platform’s own inference server: mistral.rs’s mistralrs-server. PULSE sends no prompt, no customer data, and no generated text to any third-party model API.

One server for everything

All AI services share one inference server and one model server object inside PULSE, so they share one set of agent prompts and parameters (agents/pulse.ini and the templates in agents/templates/).

The backend finds the server one of two ways:

  • Remote. When INFERENCE_REMOTE_URL is set, PULSE uses the shared inference plane at that address and starts nothing itself.
  • Local. Otherwise it starts mistralrs-server as a child process.

Speech-to-text works the same way with its own variable, VOXTRAL_REMOTE_URL, and its own mistralrs-server serving Voxtral. Text-to-speech runs inside the PULSE process (piper, through the org-runink/tts module), with no sidecar.

Both servers start in the background. If a binary or model is missing, PULSE still starts, and AI calls fail with inference server is not running instead of blocking the whole app.

Which model

INFERENCE_MODEL names the model PULSE asks the server for. When it is unset, the code default is GLM-4-9B-Chat-Q4_K_M.gguf — a leftover that mistral.rs refused to load when measured (chatglm architecture), so set a Qwen-family model explicitly; the production model set is pending the System76 benchmark. Every agent uses this one tier today. A larger tier exists in the code but routing to it was reverted.

INFERENCE_MODEL must match what the server actually loads. PULSE does not check. If they disagree, the backend asks for a model the server never loaded. The old LLAMA_* variable names are no longer read at all.

Is the model ready

Two signals report readiness, and both use the same test: the engine process is running, and a readiness check against the engine has confirmed it can serve a request.

  • GET /health returns model_ready.
  • ModelService/GetModelStatus returns ready. The app polls it on startup.

A process that has started but is still loading weights counts as not ready.

Time budgets

The model runs on CPU and one generation can take many minutes. PULSE gives each generation two clocks:

ClockDefaultOverride
Wait for the first token (queue plus prompt processing)5 minutesPULSE_FIRST_TOKEN_SECONDS
DecodeSized from the agent’s token budget and the platform decode rate, capped at 45 minutesPULSE_GENERATION_TIMEOUT_SECONDS

The decode rate comes from INFERENCE_DECODE_TOK_S, which CORE injects. If a budget is too short to produce an answer that parses, PULSE refuses up front and sends nothing to the model. The error says so and names PULSE_GENERATION_TIMEOUT_SECONDS.

Sharing the slot fairly

Optional admission control queues or refuses work when the model is busy. It is off unless PULSE_ADMIT_MAX is above zero.

VariableMeaning
PULSE_ADMIT_MAXMaximum concurrent inference calls. 0 turns admission off.
PULSE_ADMIT_QUEUEMaximum waiters before backpressure. 0 is unbounded.
PULSE_ADMIT_MODEreject (default) or degrade.
PULSE_ADMIT_WEIGHTSTier weights, for example dedicated:4,agency:2,starter:1.

When the pool is full in reject mode, the caller gets ResourceExhausted: the inference plane is at capacity and declined this request; nothing was generated — retry in a few minutes.

Repeat requests

An exact repeat of a content generation is served from a cache instead of running the model again. PULSE_DEDUP_TTL sets how long an entry lives, as a Go duration. The default is 1h.

Grounding

Generators ground their output in real sources before they write:

  • Web research. A headless Chrome scraper and extractor (internal/metasearch) reads live pages and search results. The diagnosis reads your website this way.
  • The retrieval corpus. A per-tenant vector and keyword index (store/corpus). PULSE embeds text by calling /v1/embeddings at EMBEDDING_URL. EMBEDDING_FORMAT names the embedding model’s instruction format. Unset means qwen3-embedding. The other accepted values are nomic-v1.5 and raw. An unknown value stops the server at startup.

When the embedding endpoint is unreachable, the generators still run, but they say they are ungrounded rather than invent facts. The startup log states grounding ACTIVE or grounding INACTIVE: <reason>.

Output guards

The output of content agents is screened before it reaches you. Other agents (diagnostics, radar, the assistant) stream straight through by default. To screen them too, set PULSE_EGRESS_ALL_AGENTS. It is opt-in because screening buffers the whole response, so you lose token-by-token streaming.