Sovereign models
All of FACE’s inference, including text reasoning, vision, embeddings and voice, runs on models served inside the customer’s own cluster. FACE calls no third-party LLM API. It also ships no inference binary: it is an HTTP client of the platform’s shared model plane.
The model plane contract
FACE talks to the plane through the OpenAI-compatible HTTP API:
| Purpose | Endpoint FACE calls | Base URL |
|---|---|---|
| Text and vision generation | POST /v1/chat/completions | INFERENCE_REMOTE_URL |
| Embeddings | POST /v1/embeddings | EMBEDDING_URL, falling back to INFERENCE_REMOTE_URL |
INFERENCE_REMOTE_URLis the only variable FACE reads for the plane. The vision and voice tiers (VISION_REMOTE_URL,VOICE_REMOTE_URL) default to it.INFERENCE_API_KEYis sent as a bearer token when set. FACE never generates a service token. A plane that requires one and a FACE configured without one fails per call with the plane’s own authorisation error.- Admission control is available.
FACE_ADMIT_MAXbounds concurrent inference calls at the one door every call passes through. It’s off when unset or0.
Deploying and sizing the plane is the platform’s job. See Operations.
When the plane is not configured
There is no local fallback tier, so a FACE with no plane has no inference at all. FACE refuses to run in that state, in two places:
At startup. If
INFERENCE_REMOTE_URLis unset or blank, the server refuses to start. The error names the variable and explains that FACE has no local model server.Per call. If a model call still reaches an empty base URL,
ModelServicereturns gRPCFailedPreconditionwith a message that names the variable. For example, fromEmbedding:embeddings is unavailable: no inference endpoint configured (INFERENCE_REMOTE_URL is unset or blank). FACE has no local model server; this is a deployment fault in CORE’s manifests, not a transient error
FailedPrecondition was chosen because retrying can’t fix it. It’s a configuration
state, and the cockpit can show the sentence verbatim.
ModelService
ModelService is FACE’s gRPC door to the plane. Every agent, deriver and extractor that
needs a model goes through it.
| RPC | Purpose |
|---|---|
Generate | Text or multimodal generation. |
Embedding | One embedding vector for one text. |
SynthesizeSpeech | Text to WAV audio through the in-cluster text-to-speech service. voice defaults to tara. |
SpeakFetchSummary | The spoken executive summary of the latest fetch, composed deterministically on the server. |
Task routing
GenerateRequest.task_type says which kind of model a call needs:
TaskType | Model id FACE sends | Routed to |
|---|---|---|
TASK_TYPE_GENERAL / TASK_TYPE_UNSPECIFIED | default | The general reasoning tier |
TASK_TYPE_CODE | coder | The code tier |
TASK_TYPE_VISION | vision | The vision tier |
The platform’s model router chooses the tier from that id. GenerateRequest.model is
deprecated and the server ignores it.
Images are typed
A picture travels in ImagePart (raw bytes, a mime_type, and a text_offset into
the prompt), never as base64 inside prompt text. The receiver base64-encodes it once,
when it builds the data: URL the plane expects. Because an image has its own field, a
string in untrusted text can’t be misread as an image, and text guards never evaluate
image bytes as prose. ChatMessage.images carries images per turn in multi-turn
conversations.
Embeddings
Embeddings come from a dedicated embedding model, not the chat model, because a chat
model serving /v1/embeddings returns degenerate vectors. FACE applies the model’s
input contract before embedding: a query instruction on the query side, nothing on the
document side.
EMBEDDING_FORMATselects that contract. The default isqwen3-embedding, the format of the Qwen3-Embedding-0.6B model. An unknown name stops the server at startup with the list of known names. FACE never silently falls back to raw text.MODEL_EMBEDDING_DIM(default768) is the index dimensionality. The default format is Matryoshka-trained, so its 1024-dimension vector is cut to 768 on the client. A format that isn’t Matryoshka-trained refuses the cut.- Changing the format starts a new index. The corpus index is named after the format’s vector space, so a change opens a new, empty index and leaves the old one untouched as the rollback.
- Startup probe. At boot, FACE embeds a test query through the same path a real query takes. The log says whether semantic grounding is active or inactive. See Grounding.
What this does not establish
- A configured URL is not a reachable plane. The startup refusal checks that someone wrote the variable down. It dials nothing.
- A reachable plane is not a loaded model. The plane can answer on its port while weights are still loading.
- FACE can’t verify which embedding model the endpoint serves. Nothing asks the
endpoint which model it serves. If
EMBEDDING_FORMATdisagrees with the model behindEMBEDDING_URL, the vectors have the right length and the wrong instruction, and nothing fails. - Being sovereign doesn’t make a model accurate. This page describes where inference runs, not how good its answers are. The structural checks that bound model output are on Structured output and Agents & oversight.