Skip to content

Upgrades and CD

The images

FACE ships two images. Both are built from source by CORE’s reusable image workflow (core-images-build.yml), which builds with werf and pushes to the cluster’s node-local registry.

ImageBuilt fromContents
face-backendgrpc/Dockerfile/app/server (the runink CLI; serve starts the backend), /app/entrypoint, Litestream 0.5.17 (SHA-256 pinned) with /app/litestream.yml, and a headless Chromium for rendering. It runs as a non-root user (uid 1000) on a pinned Debian slim base.
face-frontendflutter/sleeve/DockerfileThe release Flutter web bundle (flutter build web --release, with no backend host baked in) and the static sleeve, on a distroless base.

The backend build clones the shared library repositories (inference, mesh, security, store, ml, web, ui, billing) at GIT_REF (default main), generates the protobuf code, and compiles with CGO enabled. It needs a read token for those repositories, passed as the build secret gh_token. Protobuf stubs are not in the repository, so the image build always generates them.

Continuous integration

WorkflowRuns onEnforces
backend-build.ymlchanges under grpc/** (and the demo/** tree its tests read)buf generate, go build, go vet, go test. go test -race also runs weekly.
ci.yml, job flutter-analyzeevery pull request and every push to mainflutter analyze --no-fatal-infos, then flutter test. The job fails if the Dart or test file counts fall below their floors.

Continuous delivery

A push to main that touches the backend (grpc/**, go.mod, go.sum) runs cd-backend.yml. A push that touches flutter/** runs cd-frontend.yml. You can also start either workflow by hand. Both follow the same steps:

Build

The image is built with tag latest.

Preflight

The workflow checks that kubectl reaches the cluster (kubectl get --raw=/readyz). If it does not, the workflow stops and rolls nothing.

Roll

The workflow records the running image digest, deletes the pod, and waits up to 300 seconds for rollout status. The backend workflow rolls the control plane (face-control-plane) and the static face-runner. The frontend workflow rolls face-frontend.

Confirm

The workflow compares the new digest with the old one. If they match, it warns: either the build did not reach the registry, or the commit produced an identical image.

Rollouts through the operator use the Recreate strategy. The old pod stops and releases its storage lock before the new pod starts. Expect a few seconds of unavailability per roll. Do not change the strategy to rolling: two pods would contend for the same single-writer state.

Image references and pull policy

The operator pulls with Always for a tag, and with IfNotPresent for a digest (image@sha256:…). CORE_IMAGE_PULL_POLICY on the operator overrides both. For production and air-gapped sites, pin the image field of the ControlPlane and FaceInstance resources to a digest, so that every roll is reproducible and a rollback is exact. The operator reports a mutable tag on a FaceInstance through its ImageMutable condition.

When the core-mesh-ca Secret changes, the operator writes its hash into the pod template annotation core.runink.org/mesh-ca-hash. The change rolls the pods, so each one picks up the new CA.

Upgrade checklist

Read the release’s new requirements

If a release adds a boot requirement, the deployment side must land first: Prerequisites and ordering explains why. Check the namespace for core-mesh-ca, core-envelope-kek and face-session, and check that INFERENCE_REMOTE_URL is set on both the control plane and the runners.

Remove names the new image refuses

Take OBJECTSTORE_USE_SSL out of any env list. It is refused at startup. The full list of retired names is at the end of the Configuration reference.

Escrow the KEK

Confirm that the KEK is in escrow before you upgrade. Storage migrations run on first boot. See Persistence and backup.

Roll the control plane, then the runners

The control plane and runners must be able to verify each other’s mesh certificates, and they must agree on the tenant.

Verify

Check that /readyz returns ready, that the grounding line reads ACTIVE, and that a ModelService/Generate call answers. See Health and troubleshooting.

Rolling back

Storage migrations leave the old data in place: <file>.migrated files, older object-store tables and the old knowledge index. See Migrations leave a rollback in place. So an older image can usually boot against the same stores, without anything written after the upgrade. Roll back by setting the image field back to the previous digest. Do not roll back the KEK: the same KEK must stay in place.