Model plane
All of FACE’s reasoning runs on a model plane inside your cluster: text generation, embeddings, vision grading, and every agent (rules, posture, twins, compliance, the fetch cascade). FACE is an HTTP client of that plane. It ships no inference binary, starts no model process and has no local fallback. CORE runs the plane. mistral.rs is the intended engine, but nothing in FACE’s client is specific to mistral.rs.
The contract
The plane must speak the OpenAI-compatible HTTP API:
| Call | Request | Notes |
|---|---|---|
| Chat | POST <INFERENCE_REMOTE_URL>/v1/chat/completions | Always sent with "stream": true. FACE reads server-sent events. |
| Embeddings | POST <EMBEDDING_URL or INFERENCE_REMOTE_URL>/v1/embeddings | One input string per request. |
| Auth | Authorization: Bearer <INFERENCE_API_KEY> | Sent only when INFERENCE_API_KEY is non-empty. |
The model field carries a task id, not a file name: default for general reasoning,
coder for code tasks and vision for vision tasks. Embedding requests send default.
A router in front of the engines can use this id to send each task to the right tier.
Vision and voice calls go to VISION_REMOTE_URL and VOICE_REMOTE_URL. Both default to
INFERENCE_REMOTE_URL.
Bringing a plane up and pointing FACE at it
Start an OpenAI-compatible server
In a cluster, CORE’s inference manifests provide the plane. Use the Service URL CORE
gives you. For a local run, any OpenAI-compatible server works. Serve a chat model on one
port and an embedding model on another, because a chat model answering
/v1/embeddings returns unusable vectors.
Set the endpoint variables
export INFERENCE_REMOTE_URL='http://127.0.0.1:<chat-port>'
export EMBEDDING_URL='http://127.0.0.1:<embed-port>'
export MODEL_EMBEDDING_DIM=768
export EMBEDDING_FORMAT=<format matching the embedding model>In a cluster, set the same names in the resource’s env. The operator already sets them
for the ControlPlane and FaceInstance pods.
Check the boot log
A healthy start logs:
🧭 Routing reasoning to the shared inference plane at <url> (no local model server)
🧭 Embedding: tier <EMBEDDING_URL>, vector space <space>About 15 seconds later, and after up to 12 retries while the model loads, one grounding line follows (see below).
Prove the plane answers through FACE
The boot check only proves that a URL is set. To prove that a model answers, sign in and
call ModelService/Generate through FACE itself:
TOK=$(grpcurl -plaintext -d '{"method":"AUTH_METHOD_CREDENTIALS","username":"'"$ADMIN_EMAIL"'","password":"'"$ADMIN_PASSWORD"'"}' \
127.0.0.1:$PORT semantics.v1.IdentityService/Login | grep -o '"token": "[^"]*"' | cut -d'"' -f4)
test -n "$TOK" || echo "LOGIN RETURNED NO TOKEN"
grpcurl -plaintext -H "authorization: Bearer $TOK" \
-d '{"prompt":"Say hello in one sentence.","temperature":0,"max_tokens":40}' \
127.0.0.1:$PORT semantics.v1.ModelService/GenerateA reply with model text and "finishReason": "stop" means the whole path works. In a
cluster, run this against a kubectl port-forward to the pod’s app port. The login field
is token. If that field comes back empty, the next call fails with
invalid or expired session token, and the error points at the wrong call.
What each failure looks like
| Situation | What you see |
|---|---|
INFERENCE_REMOTE_URL unset | The process exits at boot with no inference endpoint configured: INFERENCE_REMOTE_URL is unset or blank. |
| An endpoint was never configured for one call path | That RPC fails with FailedPrecondition: <what> is unavailable: no inference endpoint configured (INFERENCE_REMOTE_URL is unset or blank)… |
The plane requires a token and INFERENCE_API_KEY is empty | An HTTP 401 on each call. It is not a boot failure. |
| The plane is reachable but the model is still loading, or the model id is not served | A per-call error from the plane. The boot check does not catch it. |
The embedding dimension check
The knowledge index is built for vectors of width MODEL_EMBEDDING_DIM (default 768).
EMBEDDING_FORMAT says how to wrap queries and documents for the embedding model, and how
to fit its output to that width. The default, qwen3-embedding, cuts a 1024-dimension
vector to 768. At boot, FACE embeds one probe string and logs exactly one of these lines:
✅ Comet grounding ACTIVE: embedding endpoint returns 768-dim vectors in space <space>
⚠️ Comet grounding INACTIVE: the embedding tier did not answer after warmup (<error>) — grounding falls back to DCI grep
⚠️ Comet grounding INACTIVE: the embedding endpoint returned an empty or all-zero vector (a chat model answering /v1/embeddings does this) — grounding falls back to DCI grep
⚠️ Comet grounding INACTIVE: embedding endpoint returns <n>-dim vectors, index expects <m> (MODEL_EMBEDDING_DIM), and the format <space> cannot fit them — grounding falls back to DCI grepWhen EMBEDDING_URL is unset, the tier line says so, and INACTIVE is expected:
🧭 Embedding: no dedicated embedding tier (EMBEDDING_URL unset): embeddings go to the chat plane, which serves no embedding model, so semantic grounding is expected to be INACTIVE; vector space <space>INACTIVE does not stop the process. Search falls back to keyword grep over the
document corpus, so answers still come back, but with weaker grounding. Fix it by
pointing EMBEDDING_URL at a real embedding model, and by setting MODEL_EMBEDDING_DIM
and EMBEDDING_FORMAT to match that model.
ACTIVE proves only the width of the vectors. A different model that returns the same
width looks identical from FACE. Each format and width gets its own knowledge index
(knowledge-<format>-d<dim>; raw keeps the original knowledge index). Changing
EMBEDDING_FORMAT therefore starts an empty index and leaves the old one untouched as a
rollback. FACE does not re-embed existing documents at boot.