Architecture
Platform dependencies
grpc/go.mod consumes the extracted Runink platform libraries through
relative replace directives, so a bare go build needs them checked out as
siblings of the repo:
| Library | Used for |
|---|---|
inference | The client for the shared mistral.rs plane, Voxtral STT, the admission scheduler and backpressure gate, and the judging ladder (inference/judgement) |
mesh | Interceptors (recovery, rate limit, logging), telemetry and the shared step-code vocabulary for reasoning beats |
security | OIDC verification, bearer extraction, the envelope KEK, and the outbound transport rule |
store | The object store, the appfs encrypted root, and the corpus and vector index for memory |
ui | Only the runink.ui.judgement.v1 Go types embedded in luna.pb.go (billing is in the module graph only because ui requires it) |
ml is deliberately not used. LUNA never depends on CORE, and a test
fails the build if CORE appears in the module graph or in an import.
Services
serve registers nine gRPC services on one server:
| Service | File | Role |
|---|---|---|
IdentityService | identity_server.go | Google sign-in against the allowlist, WhoAmI, Logout |
AvatarService | avatar_server.go | Painted forms (builtin_kind) or uploaded video loops. Selection is an append-only record. |
ChatService | chat_server.go, chat_actions.go | Streaming chat turns, deterministic plan actions, history |
VoiceService | voice_server.go, tts_http.go | Voxtral STT (streaming) and ttsd TTS (unary) |
VisionService | vision_server.go | Meal-photo macros and form review, sent over HTTP to the dedicated vision tier |
TrainingService | training_server.go, setlog_server.go | Plans, catalogue search, set logging, plan-from-photos |
NutritionService | nutrition_server.go, pantry_server.go | Meal plans derived from the training plan, and the pantry |
HabitService | habit_server.go, activity_server.go | Habit events, score, streaks, levels, traits, skills, activities |
ProfileService | profile_server.go | Profile, body composition, wearable days, inner context |
gRPC reflection is registered. See the API reference.
Request pipeline
One listener is split by cmux. Native gRPC (HTTP/2 application/grpc)
goes to the gRPC server. Everything else goes to an HTTP handler that serves
gRPC-Web, /media/ (signed avatar media), and the static Flutter bundle, all
from one origin.
The interceptor chain, in order:
- Recovery.
- Rate limit, keyed per session rather than per peer IP, because the gateway in front makes every peer IP identical. The defaults are 5 rps with a burst of 20.
- Logging. Conversation text is never logged.
- Auth. This validates the session, re-checks the allowlist, and attaches the account to the context so that the data layer scopes every read and write by it.
- Serialized send (streams only, innermost). Long RPCs send heartbeats
from a goroutine while results are also being sent, so every
Sendis serialized here. New streaming RPCs get this for free.
The maximum receive size is 17 MiB, derived from the payload caps. Server keepalive is a 30 s ping with a 20 s timeout, and pings without a stream are allowed. Keepalive only protects native HTTP/2 clients. For gRPC-Web, the application heartbeat is the only liveness signal.
The AI harness (internal/harness)
These are the controls enforced in Go, as opposed to prompt guidance:
- Context budget. The system prompt (persona catalogue, guidance, guardrails, active plan, targets, habits, recalled memory) is measured against the model’s window. History is trimmed oldest-first, so the engine never silently truncates the opening persona tag or a plan’s closing brace.
- Untrusted content. Photos, transcripts and pasted text are neutralised,
capped and fenced. The
[[persona:…]]syntax is defused so content can’t spoof which aspect is speaking. Vision output is scanned too, because it is journaled and replayed into later prompts. - Caps. There is a transport limiter, plus hourly model-call caps: chat 30, vision 12, voice 120.
- Heartbeat. Every RPC that waits on the model streams, and resends a
beat every 15 s until the first real output. Otherwise a silent response is
killed by the network path, which surfaces as
Request failed with status: 0. - Vision image count. At most 3 images per multimodal request, because the vision tier’s attention cost grows quadratically with patch count. The app batches beyond this limit. The server and app constants live in two languages with no shared source.
internal/openbias runs before any model call, and it can block. Its
rules have severities and cover prompt injection, credential extraction, hate
speech, self-harm, performance-enhancing drug protocols, training through
injury, disordered eating, extreme restriction, medication dosing, and animal
harm. It is wired for chat, plan-draft and meal-suggest. Its
false-positive set is tested hardest: “my shoulder hurts” and “how do I kill
fleas on my cat” stay allowed. agents/openbias_rules/*.md is a
documentation mirror, not runtime configuration.
Per-function sampling lives in agents/luna.ini (sections [*], [chat],
[plan-draft], [meal-suggest]). The sections are functions, not
aspects, because the aspect is chosen inside the generation.
Personas: routing without a second call
The system prompt asks the model to open its reply with [[persona:<id>]].
TagScanner strips the tag from the stream before the user sees it, and the
first chunk that carries text reports persona. An unknown or missing tag
falls back to counsel. persona_source tells the client whether the aspect
was model self-report or chosen by the backend (deterministic
actions, refusals, the fixed vision aspects). A reasoning beat may only be
tinted by a backend persona.
persona.ScopeForMessage decides, before the model runs, which data scopes
are assembled (Habits, Plans, Memory, Inner). An aspect that doesn’t
declare Habits can never cause the health journal to be loaded. The ids
must match the client mirror in flutter/lib/features/chat/personas.dart.
Only the Go side is tested.
Deterministic chat actions
chat_actions.go recognises plan-creation intents with regular expressions:
a create verb, plus training/workout/gym… or meal/nutrition/diet…,
plus a plan word. It then runs the same grounded generators the plan screen
uses, with no model in the loop. The reply starts in milliseconds, and it
carries a draft_plan for the client to render as an approval deck. Every
other turn goes to the model.
Journal and memory
- Every turn is journaled before the
donechunk. The write is bounded at 5 s on a detached context, sodonecan carrynot_persistedtruthfully. The reply streams either way. - Memory is one appfs vector index per account. It uses hybrid vector and BM25 search, sealed snapshots plus a WAL, and int8 storage, keyed by a stable per-turn id so upserts are idempotent. A restart recovers it without re-embedding.
- Queries and stored turns are formatted per
EMBEDDING_FORMAT. With the default Qwen3 format, recall queries carry anInstruct:prefix and stored turns carry nothing.
The judging ladder
internal/judgegate wraps inference/judgement and runs over every model
proposal a person might act on: RevisePlan, PlanFromImages, the foods in
DraftMealPlan, AnalyzeMeal and ReviewWorkoutForm. Verdicts are
CONCUR, DISSENT, UNABLE_TO_JUDGE and OUT_OF_SCOPE. Each one is
journaled in the account’s own log. Deterministic output and free chat are
not judged. LUNA_JUDGEMENT=off switches the ladder off, and that is never
treated as a pass.
Grounded generators
- Training (
internal/training,internal/exercises): an embedded catalogue of 868 movements from free-exercise-db. A test pins the count. The model only picks ids from a shortlist, so invented ids are dropped and reported. Loads come only from logged sets (progression.go). - Nutrition (
internal/nutrition): targets are pure Go arithmetic over the training week and body mass.ScaleToTargetrescales the model’s proposed foods so that the ingredients add up exactly to each meal. - Habits (
internal/habit): the score, streaks, wallet, traits, skills and avatar stage are all pure views over the journal.