Health and troubleshooting
Health endpoints
The backend serves three unauthenticated HTTP endpoints on its app port, 7100 on the
control plane and 7102 on a runner.
| Path | Answers | Use it for |
|---|---|---|
GET /livez | Always 200 with "status": "alive", once the process has reached its serving loop. It checks no shared dependency. | Liveness. A fleet-wide failure then does not restart every pod at once. |
GET /readyz | 200 ready or 503 not_ready, with one entry per check. | Readiness. When it fails, the pod is taken out of rotation but stays alive to be inspected. |
GET /auth/config | The cockpit’s sign-in configuration. | The operator’s default probe path. It answers once the web handler is serving. |
/readyz evaluates two checks:
mesh_identityfails when the rotating mesh leaf is unhealthy, or does not chain to the loaded CA. When it passes, it reports the CA’s expiry date. Watch that date: an expired CA stops every mesh connection.raft_consensusis evaluated only on a consensus member (RAFT_NODE_IDset). It fails when the member has had no leader for longer than a 30-second grace period, or when the raft loop does not answer a status query. On a runner it isskipped.
kubectl -n <ns> port-forward deploy/face-control-plane 7100:7100 &
curl -s localhost:7100/readyz | jqKubelet probes
The operator puts startup, readiness and liveness probes on the backend container.
Tune them under spec.probes of the ControlPlane or FaceInstance:
| Field | Default |
|---|---|
path | /auth/config. It is used by all three probes, and it is always the startup probe’s path. |
readinessPath, livenessPath | path |
startupSeconds | 900. This is the boot budget before the kubelet restarts the container. Probes run every 10 seconds. |
readinessFailureThreshold, livenessFailureThreshold | 3, 6. Liveness runs every 30 seconds. |
timeoutSeconds | 5 |
disabled | false. It removes all three probes. |
To make readiness reflect the mesh and consensus, set readinessPath: /readyz and
livenessPath: /livez. Do not point the startup probe or the liveness probe at
/readyz. A shared failure, such as an expired CA, would then restart every pod and
destroy the state you need to diagnose it.
Symptoms
The pod never starts: CreateContainerConfigError
A required Secret is missing from the namespace. kubectl describe pod names it. The
usual ones are core-mesh-ca, core-envelope-kek and face-session. Create or
distribute the Secret (see Prerequisites and ordering).
The pod starts on its own at the next retry.
The pod crashloops
Read the first lines of the previous container’s log:
kubectl -n <ns> logs deploy/face-control-plane -c face-backend --previous | grep -m3 -E '❌|refusing|no inference endpoint|FACE appfs root|Durable metadata store'| Log line starts with | Cause and fix |
|---|---|
❌ refusing to serve: face has no session signing key configured (AUTH_JWT_SECRET unset) | The session key is missing. Provide face-session, or set AUTH_JWT_SECRET outside a cluster. |
❌ mesh identity unavailable, refusing to start: | The CA is missing, incomplete or expired. Check both keys and the validity dates, then redistribute with core mesh ca --namespace <ns>. Restarting does not fix an expired CA. |
no inference endpoint configured: INFERENCE_REMOTE_URL is unset or blank. | Set INFERENCE_REMOTE_URL. The retired names REMOTE_LLAMA_URL and LLAMA_REMOTE_URL are not read. |
EMBEDDING_FORMAT: corpus: unknown embedding format | Use qwen3-embedding, nomic-v1.5 or raw. |
FACE appfs root: … at-rest KEK unavailable | CORE_ENVELOPE_KEK did not resolve. Check kubectl -n <ns> get secret core-envelope-kek. |
FACE appfs root: … envelope: CORE_ENVELOPE_KEK is not valid base64 or KEK must be 32 bytes | The key is malformed. Do not generate a new one if data already exists: restore the original from escrow. |
❌ FACE appfs root could not be mounted … object store configured but unavailable | The object store is unreachable, or the credentials are wrong. objectstore: refusing to send object-store traffic in cleartext to … means the endpoint needs https://. |
objectstore: OBJECTSTORE_USE_SSL is set | Remove the variable, and put the scheme in OBJECTSTORE_ENDPOINT. |
raft orchestrator failed to start and RAFT_NODE_ID=<n> is set | Consensus state could not be opened. Often the previous pod still held its lock; the next restart usually clears it. If this pod is not meant to vote, unset RAFT_NODE_ID. |
❌ RUNNER_ENROLL_ADDR is set but RUNNER_ENROLL_TOKEN is empty | A self-hosted runner is missing its token (see Runners). |
SQLITE_PATH="…" must be an absolute path or PORT="…" is not a port number | The image entrypoint rejected its input. |
litestream restore failed: | The replica exists but could not be restored. Check MINIO_* and the replica bucket. |
Running, but panels are empty or say nothing is there
FACE never substitutes sample data. An empty panel is either an empty estate or a stated failure. Work through these in order:
- No data source. The domain map answers
connect a data source first, and posture answersno data sources to assess. Connect a source (see Data sources) and run a fetch. - The fetch never reached a runner. The fetch reports
Analysis failed: dedicated runner unreachable — fetch does not run on the control plane:. Check the runner’s readiness andMANAGED_RUNNER_ENDPOINT. - The model plane is not answering. The boot log shows the configured URL, but
calls fail. Call
ModelService/Generateas shown in Model plane. AFailedPreconditionthat namesINFERENCE_REMOTE_URLmeans one call path has no endpoint. - Schedules produce nothing. Without
SCHEDULE_EXECUTE_FETCH=true, each trigger logs📅 scheduled trigger (dry-run): … — set SCHEDULE_EXECUTE_FETCH=true to run the real fetch. - The instance picker says “Instance discovery is unavailable”. The pod could not
list
ClientInstances. Its ServiceAccount needslistonclientinstances.core.runink.org.
Sign-in fails
| Message | Cause |
|---|---|
invalid or expired session token | The session was revoked or has expired, or the bearer token was empty. On runners: the runner does not share the control plane’s appfs root. |
Password sign-in is disabled — use Google sign-in | AUTH_GOOGLE_ONLY=true. Only the break-glass admin may use a password. |
SSO is not available on this instance | SSO is configured, but AUTH_ALLOWED_EMAILS is empty. The log line is sso_login_refused_no_allowlist. |
This account is not authorized for this instance | The account is not in AUTH_ALLOWED_EMAILS. The operator re-projects the list from the ClientInstance users. The pod must restart to pick it up. |
OIDC not configured (set OIDC_ISSUER and OIDC_CLIENT_ID) | SSO was attempted with no issuer or client id set. |
| A Google token is rejected for its audience | The cockpit’s google_client_id must equal OIDC_CLIENT_ID, or set OIDC_AUDIENCE. |
| The admin cannot log in at all | Look for break-glass admin not seeded: ADMIN_EMAIL and/or ADMIN_PASSWORD unset at boot. The operator takes both from <controlplane>-admin. |
Grounding is weak: embedding mismatch
Look for the Comet grounding line about 15 seconds after boot:
⚠️ Comet grounding INACTIVE: embedding endpoint returns 1024-dim vectors, index expects 768 (MODEL_EMBEDDING_DIM), and the format raw-d768 cannot fit them — grounding falls back to DCI grepSet EMBEDDING_FORMAT to the model’s own format. Only a format that allows it cuts
longer vectors: qwen3-embedding cuts 1024 to 768. Otherwise, set MODEL_EMBEDDING_DIM
to the width the model returns. An all-zero or empty vector means
EMBEDDING_URL points at a chat model. did not answer after warmup means the embedding
tier is down or unreachable. The full list is on
Model plane.
A change “does not stick”, or a roll changes nothing
- A config change is reverted. You edited the live Deployment, and the operator
reconciled it back. Change the
ControlPlaneorFaceInstanceenvinstead. - A value in
envis ignored. Some names are appended afterenv.FACE_TENANTandRUNNER_IDalways win. Same image digest after the roll. The build did not reach the registry, or the commit produced an identical image. Pin the digest to be sure what is running.RAFT_DATA_DIR unset — raft state is MEMORY-ONLY. This is fine for local development. In a cluster, the operator sets it to/raft.