Skip to content
Persistence and backup

Persistence and backup

FACE keeps its state on self-hosted storage inside your cluster. There is no managed cloud database. Every store below is sealed at rest under keys derived from one platform key-encryption key, CORE_ENVELOPE_KEK. That makes the KEK the most important thing to back up (see Key custody).

StoreWhat it holdsWhere it lives
appfs rootSessions, accounts, connections and their credentials, application metadata, fetch results, schedules, rule policies, compute budgetsThe object-store bucket face, or an encrypted local directory
MetaDB (face.db)A small SQLite key-value store on the control planeSQLITE_PATH, replicated by Litestream
Record storeSession traces (with embedding vectors) and feedback, as self-describing AvroThe object store
Knowledge corpusOne semantic index per tenant: a write-ahead log plus sealed snapshotsThe object store, under comet/<tenant>/knowledge/, or <data dir>/knowledge/
Camera evidenceEvidence frames, when EVIDENCE_FRAME_RETENTION is setThe object store
Fetch ledgerThe fetch cascade’s ledgerFETCH_LEDGER_DIR on local disk

The appfs root

At boot, FACE mounts its own application root, and it will not serve without it. When OBJECTSTORE_ENDPOINT is set, the root is the bucket face on the object store. Otherwise, it is an encrypted local directory, FACE_APPFS_DIR. The root’s key is derived from the KEK for FACE alone, and every block is bound to its full path, so a sealed block cannot be moved to another path and still open. The boot log says which backend is in use:

🗂️  FACE appfs root online: objectstore bucket face, tenant <tenant>
🗂️  FACE appfs root online: encrypted local directory <dir>/face, tenant <tenant>

The root is split into two scopes:

  • Deployment-wide: sessions (run/sessions), accounts (var/lib/auth_users) and compute budgets.
  • Per tenant (tenants/<FACE_TENANT>/, default default-tenant): connections and their credentials, application metadata, the four fetch snapshots, schedules, remediation overlays and routes, rule policies and support calls.

A runner gets its own tables at tenants/<t>/var/lib/runners.<RUNNER_ID>.<table>. The exception is the fetch snapshots, which the runner writes and the control plane reads. That shared location is why a runner and its control plane must use the same tenant.

Set FACE_TENANT before an instance’s first boot. That boot copies any legacy data into the tenant it sees, once. If you set the tenant later, the legacy data stays under default-tenant. On a runner, keep RUNNER_ID stable. The operator uses the Deployment name. If it fell back to the pod name, every restart would start the runner on a fresh, empty set of tables.

When there is no object store, the local root must be on storage that survives a restart. The operator points FACE_APPFS_DIR at appfs/ inside the pod’s persistent connections volume. If the default location next to face.db were used instead, it would be on the /data emptyDir, and sessions and credentials would be lost on every restart.

Sessions are validated on every request against this root, with a 2-second cache. A revocation on the control plane therefore takes effect on a runner within that window.

MetaDB and Litestream

The control plane keeps a SQLite database, face.db, at SQLITE_PATH. It seals every value before the value is written, so both the local file and its replica hold ciphertext. Keys are not sealed, because the store does prefix scans on them. Never put secret material in a key. Runners open no SQLite database.

The image bundles Litestream 0.5.17, pinned by SHA-256, and the configuration /app/litestream.yml:

SettingValue
ReplicaS3-compatible, path-style, litestream/face.db in MINIO_BUCKET at MINIO_ENDPOINT
Sync interval1 second
Compaction levels2 minutes, 10 minutes, 1 hour
Snapshotsevery hour, kept for 24 hours

The /data volume is an emptyDir. On every pod start, the operator’s litestream-restore init container rebuilds face.db from the replica when a replica exists and no local file does. After that, the litestream sidecar replicates continuously. When the image runs on its own (outside the operator), its entrypoint does the same when MINIO_ENDPOINT is set:

litestream: restoring /app/grpc/data/face.db from MinIO (if a replica exists)...
litestream: replicating /app/grpc/data/face.db -> MinIO; launching server...

If MINIO_ENDPOINT is unset, it prints Litestream/MinIO not configured — running server without WAL replication. A failed restore stops the container with litestream restore failed: <error>.

Credential store

A connection’s credentials never sit in an environment variable or in the connection record. The connection record has no field a secret would fit in. When a connection is saved, secret-named fields are lifted into a credential bundle. Bundles are stored as rows of the tenant table tenants/<t>/var/lib/connections (credentials/<ref>), sealed under FACE’s derived key and bound to their path. Consensus replicates only the connection metadata snapshot, never the credentials.

Where FACE writes to local disk (the local appfs root, the knowledge corpus without an object store), directories are created 0700 and files 0600. The older layout (FACE_CONNECTIONS_DIR, with connections.json and credentials/*.json) is imported once on a persistent root, and each imported file is renamed <file>.migrated.

Migrations leave a rollback in place

Each move to a new storage layout copies the data once and leaves the old copy behind, so you can go back to an older image:

  • Legacy metadata files are imported once into their table. Each one is then replaced by an envelope-sealed <file>.migrated. To read one back:

    go run -C grpc ./cmd/faceops metalog-unseal --table <table> <file>.migrated
  • Older object-store tables (tables/face/<name>/ in the object store’s own bucket) are copied into the appfs root on first open, and only into a table that has never been written. The originals are left in place.

  • The old knowledge index (comet/index.gob) is migrated once. The marker is face/migrations/comet-index-gob.v1.json. After that the gob is never written again, so an older image still boots, minus anything indexed after the switch.

Key custody

CORE_ENVELOPE_KEK is a base64-encoded 32-byte key held in the core-envelope-kek Secret. The runink operator creates it if it is absent, and never rotates or overwrites it.

Losing the KEK loses the data. The MetaDB values, the whole appfs root (sessions, accounts, connection credentials, metadata), the sealed object-store objects and the knowledge snapshots cannot be decrypted without it. No recovery path exists, by design. Export the Secret to your offline key escrow as soon as it is created. Back it up separately from the data it protects.

To change keys, keep the old key readable. CORE_ENVELOPE_KEK_RETIRED takes a list of earlier KEKs (base64, separated by commas or whitespace). FACE and the storage libraries use them for decryption only, and seal new writes under CORE_ENVELOPE_KEK. If you replace the KEK without listing the old one as retired, everything sealed under the old key becomes unreadable.

What to back up

  1. The core-envelope-kek Secret, escrowed offline.
  2. The object-store buckets FACE writes to: the face appfs bucket, the bucket in OBJECTSTORE_BUCKET, and the Litestream replica bucket.
  3. When there is no object store: the node-local connections volume, which contains the appfs root, and the knowledge directory.
  4. The node-local /raft volume, which holds the control plane’s consensus state (RAFT_DATA_DIR).

To restore, put the KEK Secret back first, then the object-store data, and only then start the control plane. The restore init container rebuilds face.db from the replica.