Skip to content

Architecture

Platform dependencies

grpc/go.mod consumes the extracted Runink platform libraries through relative replace directives, so a bare go build needs them checked out as siblings of the repo:

LibraryUsed for
inferenceThe client for the shared mistral.rs plane, Voxtral STT, the admission scheduler and backpressure gate, and the judging ladder (inference/judgement)
meshInterceptors (recovery, rate limit, logging), telemetry and the shared step-code vocabulary for reasoning beats
securityOIDC verification, bearer extraction, the envelope KEK, and the outbound transport rule
storeThe object store, the appfs encrypted root, and the corpus and vector index for memory
uiOnly the runink.ui.judgement.v1 Go types embedded in luna.pb.go (billing is in the module graph only because ui requires it)

ml is deliberately not used. LUNA never depends on CORE, and a test fails the build if CORE appears in the module graph or in an import.

Services

serve registers nine gRPC services on one server:

ServiceFileRole
IdentityServiceidentity_server.goGoogle sign-in against the allowlist, WhoAmI, Logout
AvatarServiceavatar_server.goPainted forms (builtin_kind) or uploaded video loops. Selection is an append-only record.
ChatServicechat_server.go, chat_actions.goStreaming chat turns, deterministic plan actions, history
VoiceServicevoice_server.go, tts_http.goVoxtral STT (streaming) and ttsd TTS (unary)
VisionServicevision_server.goMeal-photo macros and form review, sent over HTTP to the dedicated vision tier
TrainingServicetraining_server.go, setlog_server.goPlans, catalogue search, set logging, plan-from-photos
NutritionServicenutrition_server.go, pantry_server.goMeal plans derived from the training plan, and the pantry
HabitServicehabit_server.go, activity_server.goHabit events, score, streaks, levels, traits, skills, activities
ProfileServiceprofile_server.goProfile, body composition, wearable days, inner context

gRPC reflection is registered. See the API reference.

Request pipeline

One listener is split by cmux. Native gRPC (HTTP/2 application/grpc) goes to the gRPC server. Everything else goes to an HTTP handler that serves gRPC-Web, /media/ (signed avatar media), and the static Flutter bundle, all from one origin.

The interceptor chain, in order:

  1. Recovery.
  2. Rate limit, keyed per session rather than per peer IP, because the gateway in front makes every peer IP identical. The defaults are 5 rps with a burst of 20.
  3. Logging. Conversation text is never logged.
  4. Auth. This validates the session, re-checks the allowlist, and attaches the account to the context so that the data layer scopes every read and write by it.
  5. Serialized send (streams only, innermost). Long RPCs send heartbeats from a goroutine while results are also being sent, so every Send is serialized here. New streaming RPCs get this for free.

The maximum receive size is 17 MiB, derived from the payload caps. Server keepalive is a 30 s ping with a 20 s timeout, and pings without a stream are allowed. Keepalive only protects native HTTP/2 clients. For gRPC-Web, the application heartbeat is the only liveness signal.

The AI harness (internal/harness)

These are the controls enforced in Go, as opposed to prompt guidance:

  • Context budget. The system prompt (persona catalogue, guidance, guardrails, active plan, targets, habits, recalled memory) is measured against the model’s window. History is trimmed oldest-first, so the engine never silently truncates the opening persona tag or a plan’s closing brace.
  • Untrusted content. Photos, transcripts and pasted text are neutralised, capped and fenced. The [[persona:…]] syntax is defused so content can’t spoof which aspect is speaking. Vision output is scanned too, because it is journaled and replayed into later prompts.
  • Caps. There is a transport limiter, plus hourly model-call caps: chat 30, vision 12, voice 120.
  • Heartbeat. Every RPC that waits on the model streams, and resends a beat every 15 s until the first real output. Otherwise a silent response is killed by the network path, which surfaces as Request failed with status: 0.
  • Vision image count. At most 3 images per multimodal request, because the vision tier’s attention cost grows quadratically with patch count. The app batches beyond this limit. The server and app constants live in two languages with no shared source.

internal/openbias runs before any model call, and it can block. Its rules have severities and cover prompt injection, credential extraction, hate speech, self-harm, performance-enhancing drug protocols, training through injury, disordered eating, extreme restriction, medication dosing, and animal harm. It is wired for chat, plan-draft and meal-suggest. Its false-positive set is tested hardest: “my shoulder hurts” and “how do I kill fleas on my cat” stay allowed. agents/openbias_rules/*.md is a documentation mirror, not runtime configuration.

Per-function sampling lives in agents/luna.ini (sections [*], [chat], [plan-draft], [meal-suggest]). The sections are functions, not aspects, because the aspect is chosen inside the generation.

Personas: routing without a second call

The system prompt asks the model to open its reply with [[persona:<id>]]. TagScanner strips the tag from the stream before the user sees it, and the first chunk that carries text reports persona. An unknown or missing tag falls back to counsel. persona_source tells the client whether the aspect was model self-report or chosen by the backend (deterministic actions, refusals, the fixed vision aspects). A reasoning beat may only be tinted by a backend persona.

persona.ScopeForMessage decides, before the model runs, which data scopes are assembled (Habits, Plans, Memory, Inner). An aspect that doesn’t declare Habits can never cause the health journal to be loaded. The ids must match the client mirror in flutter/lib/features/chat/personas.dart. Only the Go side is tested.

Deterministic chat actions

chat_actions.go recognises plan-creation intents with regular expressions: a create verb, plus training/workout/gym… or meal/nutrition/diet…, plus a plan word. It then runs the same grounded generators the plan screen uses, with no model in the loop. The reply starts in milliseconds, and it carries a draft_plan for the client to render as an approval deck. Every other turn goes to the model.

Journal and memory

  • Every turn is journaled before the done chunk. The write is bounded at 5 s on a detached context, so done can carry not_persisted truthfully. The reply streams either way.
  • Memory is one appfs vector index per account. It uses hybrid vector and BM25 search, sealed snapshots plus a WAL, and int8 storage, keyed by a stable per-turn id so upserts are idempotent. A restart recovers it without re-embedding.
  • Queries and stored turns are formatted per EMBEDDING_FORMAT. With the default Qwen3 format, recall queries carry an Instruct: prefix and stored turns carry nothing.

The judging ladder

internal/judgegate wraps inference/judgement and runs over every model proposal a person might act on: RevisePlan, PlanFromImages, the foods in DraftMealPlan, AnalyzeMeal and ReviewWorkoutForm. Verdicts are CONCUR, DISSENT, UNABLE_TO_JUDGE and OUT_OF_SCOPE. Each one is journaled in the account’s own log. Deterministic output and free chat are not judged. LUNA_JUDGEMENT=off switches the ladder off, and that is never treated as a pass.

Grounded generators

  • Training (internal/training, internal/exercises): an embedded catalogue of 868 movements from free-exercise-db. A test pins the count. The model only picks ids from a shortlist, so invented ids are dropped and reported. Loads come only from logged sets (progression.go).
  • Nutrition (internal/nutrition): targets are pure Go arithmetic over the training week and body mass. ScaleToTarget rescales the model’s proposed foods so that the ingredients add up exactly to each meal.
  • Habits (internal/habit): the score, streaks, wallet, traits, skills and avatar stage are all pure views over the journal.