Upgrades and CD
The images
FACE ships two images. Both are built from source by CORE’s reusable image workflow
(core-images-build.yml), which builds with werf and pushes to the cluster’s
node-local registry.
| Image | Built from | Contents |
|---|---|---|
face-backend | grpc/Dockerfile | /app/server (the runink CLI; serve starts the backend), /app/entrypoint, Litestream 0.5.17 (SHA-256 pinned) with /app/litestream.yml, and a headless Chromium for rendering. It runs as a non-root user (uid 1000) on a pinned Debian slim base. |
face-frontend | flutter/sleeve/Dockerfile | The release Flutter web bundle (flutter build web --release, with no backend host baked in) and the static sleeve, on a distroless base. |
The backend build clones the shared library repositories (inference, mesh,
security, store, ml, web, ui, billing) at GIT_REF (default main),
generates the protobuf code, and compiles with CGO enabled. It needs a read token for
those repositories, passed as the build secret gh_token. Protobuf stubs are not in the
repository, so the image build always generates them.
Continuous integration
| Workflow | Runs on | Enforces |
|---|---|---|
backend-build.yml | changes under grpc/** (and the demo/** tree its tests read) | buf generate, go build, go vet, go test. go test -race also runs weekly. |
ci.yml, job flutter-analyze | every pull request and every push to main | flutter analyze --no-fatal-infos, then flutter test. The job fails if the Dart or test file counts fall below their floors. |
Continuous delivery
A push to main that touches the backend (grpc/**, go.mod, go.sum) runs
cd-backend.yml. A push that touches flutter/** runs cd-frontend.yml. You can also
start either workflow by hand. Both follow the same steps:
Build
The image is built with tag latest.
Preflight
The workflow checks that kubectl reaches the cluster (kubectl get --raw=/readyz). If
it does not, the workflow stops and rolls nothing.
Roll
The workflow records the running image digest, deletes the pod, and waits up to 300
seconds for rollout status. The backend workflow rolls the control plane
(face-control-plane) and the static face-runner. The frontend workflow rolls
face-frontend.
Confirm
The workflow compares the new digest with the old one. If they match, it warns: either the build did not reach the registry, or the commit produced an identical image.
Rollouts through the operator use the Recreate strategy. The old pod stops and releases
its storage lock before the new pod starts. Expect a few seconds of unavailability per
roll. Do not change the strategy to rolling: two pods would contend for the same
single-writer state.
Image references and pull policy
The operator pulls with Always for a tag, and with IfNotPresent for a digest
(image@sha256:…). CORE_IMAGE_PULL_POLICY on the operator overrides both. For
production and air-gapped sites, pin the image field of the ControlPlane and
FaceInstance resources to a digest, so that every roll is reproducible and a rollback
is exact. The operator reports a mutable tag on a FaceInstance through its ImageMutable
condition.
When the core-mesh-ca Secret changes, the operator writes its hash into the pod
template annotation core.runink.org/mesh-ca-hash. The change rolls the pods, so each one
picks up the new CA.
Upgrade checklist
Read the release’s new requirements
If a release adds a boot requirement, the deployment side must land first:
Prerequisites and ordering explains why. Check the
namespace for core-mesh-ca, core-envelope-kek and face-session, and check that
INFERENCE_REMOTE_URL is set on both the control plane and the runners.
Remove names the new image refuses
Take OBJECTSTORE_USE_SSL out of any env list. It is refused at startup. The full list
of retired names is at the end of the
Configuration reference.
Escrow the KEK
Confirm that the KEK is in escrow before you upgrade. Storage migrations run on first boot. See Persistence and backup.
Roll the control plane, then the runners
The control plane and runners must be able to verify each other’s mesh certificates, and they must agree on the tenant.
Verify
Check that /readyz returns ready, that the grounding line reads ACTIVE, and that a
ModelService/Generate call answers. See
Health and troubleshooting.
Rolling back
Storage migrations leave the old data in place: <file>.migrated files, older
object-store tables and the old knowledge index. See
Migrations leave a rollback in place.
So an older image can usually boot against the same stores, without anything written
after the upgrade. Roll back by setting the image field back to the previous digest.
Do not roll back the KEK: the same KEK must stay in place.