Skip to content

Sovereign inference

FACE calls no third-party LLM API. Every model FACE uses for text, vision and speech runs inside your own cluster, on the Runink inference plane. Prompts, fetched records and model answers never go to a model vendor.

How inference is wired

  • One endpoint. All text generation, embeddings, vision grading and agent reasoning are HTTP calls to the in-cluster inference plane (a mistral.rs server). The base URL comes from INFERENCE_REMOTE_URL, the only variable FACE reads for it. The vision and voice tiers can be pointed elsewhere with VISION_REMOTE_URL and VOICE_REMOTE_URL. When those are unset, both fall back to INFERENCE_REMOTE_URL.
  • No bundled model server. FACE ships no inference binary and starts no model-server child process. It has no local “fallback” model tier.
  • Inference is required to start. If INFERENCE_REMOTE_URL is unset or blank, the backend refuses to start with no inference endpoint configured: INFERENCE_REMOTE_URL is unset or blank. If a call reaches a model client with no endpoint, it fails with FailedPrecondition naming the variable, rather than a transport error that names nothing.
  • An authenticated plane. FACE presents the service token from INFERENCE_API_KEY to the plane. It never generates a token of its own.
  • Internal-only model service. FACE’s own ModelService gRPC door, which its agents use to reach the plane, accepts loopback callers only. See Identity & access.
  • Speech. Phone-call transcription is a request to the plane’s voice tier. Speech synthesis runs in process with a bundled local engine (Piper), so generating audio does not leave the pod.
  • No LLM vendor SDK. FACE’s Go module does not depend on any LLM vendor SDK, and the backend does not use langchaingo. Model output is parsed as TOON (see AI safety).

What sovereign means for data egress

Keeping inference in the cluster removes the largest egress path an AI product usually has. It does not mean FACE never talks to the outside world. These are the paths that can leave your network, and each one exists only if you configure the feature behind it:

PathWhat leavesWhen it exists
Data-source connectorsQueries and API requests to the source systems you connect: warehouses, ERPs, logistics APIs, devices.For each connection you create. The destinations are governed by the connector address policy.
Web search and page readingSearch terms and page requests to the public web, fetched through a headless browser.When an agent or the cockpit uses web search or reads a web page.
Maps and routingLocation and routing requests to Google Maps and Routes.Only for connections that hold a Maps API key.
Single sign-onToken verification against your OIDC issuer.Only when OIDC_ISSUER and OIDC_CLIENT_ID are set.
EmailMessages the twins agent drafts or sends through a connected mailbox.Only with a connected mailbox. Sending without a human needs an approved-recipient list (see AI safety).
Telephony and messagingCall audio and messages carried by the telephony provider, Twilio.Only when the voice line or messaging is configured.

Everything a model reads (fetched rows, page text, call transcripts) is processed by the in-cluster plane. The table above lists the only routes by which data leaves.

What these controls do not establish

  • A configured endpoint is not a reachable plane. The startup check confirms that INFERENCE_REMOTE_URL is set. It dials nothing, and it does not prove a model is loaded.
  • Sovereignty depends on where the plane runs. FACE sends inference to the URL it is given. That the URL points at a plane inside your cluster is a deployment fact, and you have to verify it on your own installation.
  • Voice calls go through a carrier. A phone call’s audio travels through the telephony provider before it reaches FACE. Transcription and reasoning are sovereign, but carriage of the call is not.
  • The dependency check is a fact about today’s build. No LLM vendor SDK is in the module today. This page does not claim an automated gate that stops one being added.