Sovereign inference
Every AI feature in PULSE runs on the Runink platform’s own inference server:
mistral.rs’s mistralrs-server. PULSE sends no prompt, no customer data, and no
generated text to any third-party model API.
One server for everything
All AI services share one inference server and one model server object inside
PULSE, so they share one set of agent prompts and parameters
(agents/pulse.ini and the templates in agents/templates/).
The backend finds the server one of two ways:
- Remote. When
INFERENCE_REMOTE_URLis set, PULSE uses the shared inference plane at that address and starts nothing itself. - Local. Otherwise it starts
mistralrs-serveras a child process.
Speech-to-text works the same way with its own variable, VOXTRAL_REMOTE_URL,
and its own mistralrs-server serving Voxtral. Text-to-speech runs inside the
PULSE process (piper, through the org-runink/tts module), with no sidecar.
Both servers start in the background. If a binary or model is missing, PULSE
still starts, and AI calls fail with inference server is not running instead
of blocking the whole app.
Which model
INFERENCE_MODEL names the model PULSE asks the server for. When it is unset,
the code default is GLM-4-9B-Chat-Q4_K_M.gguf — a leftover that mistral.rs refused to
load when measured (chatglm architecture), so set a Qwen-family model explicitly; the
production model set is pending the System76 benchmark. Every agent uses this one tier
today. A larger tier exists in the code but routing to it was reverted.
INFERENCE_MODEL must match what the server actually loads. PULSE does not
check. If they disagree, the backend asks for a model the server never loaded.
The old LLAMA_* variable names are no longer read at all.Is the model ready
Two signals report readiness, and both use the same test: the engine process is running, and a readiness check against the engine has confirmed it can serve a request.
GET /healthreturnsmodel_ready.ModelService/GetModelStatusreturnsready. The app polls it on startup.
A process that has started but is still loading weights counts as not ready.
Time budgets
The model runs on CPU and one generation can take many minutes. PULSE gives each generation two clocks:
| Clock | Default | Override |
|---|---|---|
| Wait for the first token (queue plus prompt processing) | 5 minutes | PULSE_FIRST_TOKEN_SECONDS |
| Decode | Sized from the agent’s token budget and the platform decode rate, capped at 45 minutes | PULSE_GENERATION_TIMEOUT_SECONDS |
The decode rate comes from INFERENCE_DECODE_TOK_S, which CORE injects. If a
budget is too short to produce an answer that parses, PULSE refuses up front
and sends nothing to the model. The error says so and names
PULSE_GENERATION_TIMEOUT_SECONDS.
Sharing the slot fairly
Optional admission control queues or refuses work when the model is busy. It is
off unless PULSE_ADMIT_MAX is above zero.
| Variable | Meaning |
|---|---|
PULSE_ADMIT_MAX | Maximum concurrent inference calls. 0 turns admission off. |
PULSE_ADMIT_QUEUE | Maximum waiters before backpressure. 0 is unbounded. |
PULSE_ADMIT_MODE | reject (default) or degrade. |
PULSE_ADMIT_WEIGHTS | Tier weights, for example dedicated:4,agency:2,starter:1. |
When the pool is full in reject mode, the caller gets ResourceExhausted:
the inference plane is at capacity and declined this request; nothing was generated — retry in a few minutes.
Repeat requests
An exact repeat of a content generation is served from a cache instead of
running the model again. PULSE_DEDUP_TTL sets how long an entry lives, as a
Go duration. The default is 1h.
Grounding
Generators ground their output in real sources before they write:
- Web research. A headless Chrome scraper and extractor (
internal/metasearch) reads live pages and search results. The diagnosis reads your website this way. - The retrieval corpus. A per-tenant vector and keyword index
(
store/corpus). PULSE embeds text by calling/v1/embeddingsatEMBEDDING_URL.EMBEDDING_FORMATnames the embedding model’s instruction format. Unset meansqwen3-embedding. The other accepted values arenomic-v1.5andraw. An unknown value stops the server at startup.
When the embedding endpoint is unreachable, the generators still run, but they
say they are ungrounded rather than invent facts. The startup log states
grounding ACTIVE or grounding INACTIVE: <reason>.
Output guards
The output of content agents is screened before it reaches you. Other agents
(diagnostics, radar, the assistant) stream straight through by default. To
screen them too, set PULSE_EGRESS_ALL_AGENTS. It is opt-in because screening
buffers the whole response, so you lose token-by-token streaming.