Skip to content
Deployment and storage

Deployment and storage

The image

grpc/Dockerfile builds one image, luna-backend, in three stages:

  1. The Flutter web bundle (cirruslabs/flutter:3.44.0, Dart 3.12.0), built with no source maps. The build fails if a .map file appears.
  2. A static Go binary, luna-server (CGO_ENABLED=0).
  3. A slim Debian runtime running as numeric UID 1000, with LUNA_WEB_DIR=/app/web and LUNA_AGENT_CONFIG=/app/agents/luna.ini. It exposes 8080, and the entrypoint is /app/luna-server serve.

The Flutter image and the pubspec’s Dart floor must move together. A pubspec that needs a newer Dart than the image ships builds locally and fails in-cluster. A CI step compares the two.

One port serves everything: native gRPC, gRPC-Web for the browser, the /media/ proxy for signed avatar-media URLs, and the static web bundle, which shares its origin with the API. There is no /health endpoint. Readiness means the port accepts gRPC (grpcurl … list).

CI/CD

WorkflowTriggerWhat it does
ci.ymlEvery PR and every push to maingo build / vet / test in every Go module (with the private sibling libraries checked out), and flutter analyze / test with minimum-test-count ratchets
cd.ymlPush to main, or workflow_dispatchBuilds the image through CORE’s reusable core-images-build.yml, restarts luna-system/luna, waits for Ready, and warns if the running image digest didn’t change. roll_only restarts without rebuilding.

The CD gate is the image build itself, because the Dockerfile compiles Go and builds Flutter web. CI results are advisory signals: the branch has no protection rule that blocks a merge.

Pod posture

ControlStatus
Non-root, no capabilities, read-only root filesystem, no service-account tokenEnforced by the pod securityContext
Allowlist-only accessEnforced and fail-closed
Secrets from luna-secrets via secretKeyRef, never loggedEnforced
Egress allowlistNot effective today. See below.

Storage: the appfs root

Everything LUNA persists for an account lives in one encrypted, append-only store/appfs root for the app luna. It is sealed under a key derived by HKDF from CORE_ENVELOPE_KEK and kept in the configured bucket on the shared object store:

luna/etc/allowlist.json              last allowlist enforced (sessions)
luna/home/<owner>/account            hash → address note
luna/home/<owner>/migrated/<stream>  migration receipts
luna/var/log/<stream>.<owner>        chat_turns, habit_events, avatars, plans, body, training_log
luna/var/lib/comet/memory-<space>.<owner>/   long-term memory (vector index)
luna/run/sessions/                   sessions
  • <owner> is a hash of the account address, not the address itself.
  • The streams are append-only logs. There is no delete path, and an “update” is a superseding append.
  • <space> names the embedding vector space (for example qwen3-embedding-d768). Changing EMBEDDING_FORMAT or LUNA_EMBEDDING_DIM opens a new, empty index. Re-embedding existing memories is a separate platform runbook, and nothing in LUNA runs it.
  • Avatar media is not in appfs. Custom video loops live in the object store’s public media area and are served through signed, expiring URLs. No appfs key can be reached through /media/.

Legacy data

Records written before per-account scoping sit in a partition called default. They carry no identity and are readable by nobody until an operator runs claim-legacy-data --account <email>. Records from the older record-store layout are copied into appfs by migrate-appfs, which runs after every connect and again 15 minutes later. It is idempotent and never deletes the old objects.

Persistence connects lazily

The object store connects lazily, with retries, plus a background reconnect loop. After a pod restart, the new pod IP can take a minute or two to be admitted by the network policy guarding the object store. During that window dials time out (they are dropped, not refused), and the loop keeps trying until they succeed.

Until the object store connects, memory and sessions use the local encrypted fallback root. Journals are refused honestly during that time: a chat turn still streams, but it is marked not_persisted and the app shows it as “not saved”.

Network policy

The deployment manifest, which lives in CORE’s infrastructure tree, declares a luna-ingress policy and an intended egress allowlist: DNS, the inference plane, and object storage, plus 443 to external addresses for Google’s JWKS at login.

Don’t cite the egress policy as a live control. Its last rule allows 443 to every external address, and on this node an egress ports: stanza never matches, so in practice all external traffic is allowed. The protection LUNA enforces itself is the startup outbound transport check and the absence of any third-party AI code path. Check the live state with kubectl -n luna-system get netpol before relying on it.