Skip to content

Sovereign models

All of FACE’s inference, including text reasoning, vision, embeddings and voice, runs on models served inside the customer’s own cluster. FACE calls no third-party LLM API. It also ships no inference binary: it is an HTTP client of the platform’s shared model plane.

The model plane contract

FACE talks to the plane through the OpenAI-compatible HTTP API:

PurposeEndpoint FACE callsBase URL
Text and vision generationPOST /v1/chat/completionsINFERENCE_REMOTE_URL
EmbeddingsPOST /v1/embeddingsEMBEDDING_URL, falling back to INFERENCE_REMOTE_URL
  • INFERENCE_REMOTE_URL is the only variable FACE reads for the plane. The vision and voice tiers (VISION_REMOTE_URL, VOICE_REMOTE_URL) default to it.
  • INFERENCE_API_KEY is sent as a bearer token when set. FACE never generates a service token. A plane that requires one and a FACE configured without one fails per call with the plane’s own authorisation error.
  • Admission control is available. FACE_ADMIT_MAX bounds concurrent inference calls at the one door every call passes through. It’s off when unset or 0.

Deploying and sizing the plane is the platform’s job. See Operations.

When the plane is not configured

There is no local fallback tier, so a FACE with no plane has no inference at all. FACE refuses to run in that state, in two places:

  1. At startup. If INFERENCE_REMOTE_URL is unset or blank, the server refuses to start. The error names the variable and explains that FACE has no local model server.

  2. Per call. If a model call still reaches an empty base URL, ModelService returns gRPC FailedPrecondition with a message that names the variable. For example, from Embedding:

    embeddings is unavailable: no inference endpoint configured (INFERENCE_REMOTE_URL is unset or blank). FACE has no local model server; this is a deployment fault in CORE’s manifests, not a transient error

FailedPrecondition was chosen because retrying can’t fix it. It’s a configuration state, and the cockpit can show the sentence verbatim.

ModelService

ModelService is FACE’s gRPC door to the plane. Every agent, deriver and extractor that needs a model goes through it.

RPCPurpose
GenerateText or multimodal generation.
EmbeddingOne embedding vector for one text.
SynthesizeSpeechText to WAV audio through the in-cluster text-to-speech service. voice defaults to tara.
SpeakFetchSummaryThe spoken executive summary of the latest fetch, composed deterministically on the server.

Task routing

GenerateRequest.task_type says which kind of model a call needs:

TaskTypeModel id FACE sendsRouted to
TASK_TYPE_GENERAL / TASK_TYPE_UNSPECIFIEDdefaultThe general reasoning tier
TASK_TYPE_CODEcoderThe code tier
TASK_TYPE_VISIONvisionThe vision tier

The platform’s model router chooses the tier from that id. GenerateRequest.model is deprecated and the server ignores it.

Images are typed

A picture travels in ImagePart (raw bytes, a mime_type, and a text_offset into the prompt), never as base64 inside prompt text. The receiver base64-encodes it once, when it builds the data: URL the plane expects. Because an image has its own field, a string in untrusted text can’t be misread as an image, and text guards never evaluate image bytes as prose. ChatMessage.images carries images per turn in multi-turn conversations.

Embeddings

Embeddings come from a dedicated embedding model, not the chat model, because a chat model serving /v1/embeddings returns degenerate vectors. FACE applies the model’s input contract before embedding: a query instruction on the query side, nothing on the document side.

  • EMBEDDING_FORMAT selects that contract. The default is qwen3-embedding, the format of the Qwen3-Embedding-0.6B model. An unknown name stops the server at startup with the list of known names. FACE never silently falls back to raw text.
  • MODEL_EMBEDDING_DIM (default 768) is the index dimensionality. The default format is Matryoshka-trained, so its 1024-dimension vector is cut to 768 on the client. A format that isn’t Matryoshka-trained refuses the cut.
  • Changing the format starts a new index. The corpus index is named after the format’s vector space, so a change opens a new, empty index and leaves the old one untouched as the rollback.
  • Startup probe. At boot, FACE embeds a test query through the same path a real query takes. The log says whether semantic grounding is active or inactive. See Grounding.

What this does not establish

  • A configured URL is not a reachable plane. The startup refusal checks that someone wrote the variable down. It dials nothing.
  • A reachable plane is not a loaded model. The plane can answer on its port while weights are still loading.
  • FACE can’t verify which embedding model the endpoint serves. Nothing asks the endpoint which model it serves. If EMBEDDING_FORMAT disagrees with the model behind EMBEDDING_URL, the vectors have the right length and the wrong instruction, and nothing fails.
  • Being sovereign doesn’t make a model accurate. This page describes where inference runs, not how good its answers are. The structural checks that bound model output are on Structured output and Agents & oversight.