Sovereign inference
FACE calls no third-party LLM API. Every model FACE uses for text, vision and speech runs inside your own cluster, on the Runink inference plane. Prompts, fetched records and model answers never go to a model vendor.
How inference is wired
- One endpoint. All text generation, embeddings, vision grading and agent
reasoning are HTTP calls to the in-cluster inference plane (a mistral.rs
server). The base URL comes from
INFERENCE_REMOTE_URL, the only variable FACE reads for it. The vision and voice tiers can be pointed elsewhere withVISION_REMOTE_URLandVOICE_REMOTE_URL. When those are unset, both fall back toINFERENCE_REMOTE_URL. - No bundled model server. FACE ships no inference binary and starts no model-server child process. It has no local “fallback” model tier.
- Inference is required to start. If
INFERENCE_REMOTE_URLis unset or blank, the backend refuses to start withno inference endpoint configured: INFERENCE_REMOTE_URL is unset or blank. If a call reaches a model client with no endpoint, it fails withFailedPreconditionnaming the variable, rather than a transport error that names nothing. - An authenticated plane. FACE presents the service token from
INFERENCE_API_KEYto the plane. It never generates a token of its own. - Internal-only model service. FACE’s own
ModelServicegRPC door, which its agents use to reach the plane, accepts loopback callers only. See Identity & access. - Speech. Phone-call transcription is a request to the plane’s voice tier. Speech synthesis runs in process with a bundled local engine (Piper), so generating audio does not leave the pod.
- No LLM vendor SDK. FACE’s Go module does not depend on any LLM vendor SDK,
and the backend does not use
langchaingo. Model output is parsed as TOON (see AI safety).
What sovereign means for data egress
Keeping inference in the cluster removes the largest egress path an AI product usually has. It does not mean FACE never talks to the outside world. These are the paths that can leave your network, and each one exists only if you configure the feature behind it:
| Path | What leaves | When it exists |
|---|---|---|
| Data-source connectors | Queries and API requests to the source systems you connect: warehouses, ERPs, logistics APIs, devices. | For each connection you create. The destinations are governed by the connector address policy. |
| Web search and page reading | Search terms and page requests to the public web, fetched through a headless browser. | When an agent or the cockpit uses web search or reads a web page. |
| Maps and routing | Location and routing requests to Google Maps and Routes. | Only for connections that hold a Maps API key. |
| Single sign-on | Token verification against your OIDC issuer. | Only when OIDC_ISSUER and OIDC_CLIENT_ID are set. |
| Messages the twins agent drafts or sends through a connected mailbox. | Only with a connected mailbox. Sending without a human needs an approved-recipient list (see AI safety). | |
| Telephony and messaging | Call audio and messages carried by the telephony provider, Twilio. | Only when the voice line or messaging is configured. |
Everything a model reads (fetched rows, page text, call transcripts) is processed by the in-cluster plane. The table above lists the only routes by which data leaves.
What these controls do not establish
- A configured endpoint is not a reachable plane. The startup check
confirms that
INFERENCE_REMOTE_URLis set. It dials nothing, and it does not prove a model is loaded. - Sovereignty depends on where the plane runs. FACE sends inference to the URL it is given. That the URL points at a plane inside your cluster is a deployment fact, and you have to verify it on your own installation.
- Voice calls go through a carrier. A phone call’s audio travels through the telephony provider before it reaches FACE. Transcription and reasoning are sovereign, but carriage of the call is not.
- The dependency check is a fact about today’s build. No LLM vendor SDK is in the module today. This page does not claim an automated gate that stops one being added.