The planner
The model proposes; you approve. A free-form chat message goes to
ForgeService.ProposePlan. The server hands it to the platform’s own in-cluster
model and returns a ticket at once. The studio then polls GetPlan.
How a plan is made
- The model is sovereign. The planner sends a plain
net/httprequest toINFERENCE_URL/v1/chat/completionswith the model name fromINFERENCE_MODEL. There is no SDK and no external endpoint.INFERENCE_URLshould point at the general tier; the coder tier is for the coding agents. - It is slow on purpose. Planning runs at CPU decode speed. Each plan runs under its
own 12-minute context with
max_tokens500. - One plan at a time. A second
ProposePlanwhile one is thinking fails immediately withResourceExhausted:the planner is still thinking about the previous message — the platform's own model decodes one plan at a time on CPU; wait for that plan, then ask again. It is not queued. - No grammar-constrained decoding. The model is asked for compact JSON. The object is extracted (prose and code fences are tolerated) and then validated.
- Plans live in memory only. They are kept for one hour after finishing. A restart
forgets them, and
GetPlanthen answersNotFound:the planner forgot this plan (restart or expiry) — ask again. Nothing about a plan goes to GitHub or appfs, andProposePlanwrites no audit entry.
A message may be up to 4000 characters. An empty one gets say what you want the forge to do.
The answer is untrusted
The server re-validates every step. Invalid steps are dropped, never repaired, and the reply says how many were dropped:
- the action is
neworchange; - the name follows the repo-name rule; a
changemust name an existing app, and anewmust not; - the kind maps to an
AppKind; - the brief is within the same length rules as a filed brief;
- there are at most 5 steps.
Errors from the inference plane reach Plan.error unparaphrased.
Every step is judged before you see it
FORGE mounts the shared Runink plan gate (inference/judgement JudgePlan). The verdict
goes on ProposedStep.judgement (runink.ui.judgement.v1.Judgement). The evidence is your
message, the organisation’s forged apps (when GitHub could be read) and the selected app.
The claim is always “the owner asked for this step”, never the model’s own rationale.
| Verdict | What the studio shows |
|---|---|
| Concur | The step, as a normal proposal |
| Dissent | Marked “judge dissents” with the reason, accepted = false, offered only as Use anyway |
| Unable to judge | Marked as such. Never shown as agreement |
| Unset | Not judged (FORGE_JUDGEMENT=off) |
Any verdict the model wrote itself is stripped before judging, and the reply says how many. The brief itself is never altered: it reaches the agent verbatim.
Approval
Whether it was judged or not, a proposed step is only a prefill. Choosing it puts a
typed command in the console input, for example new pipeline <name>: <brief> or
change <name>: <brief>. You may edit it, and nothing happens until you send it. The
send goes through the same deterministic command path as anything you type, and the
audit records it as yours.
Sources
grpc/internal/planner/{planner,model,judge}.go; grpc/cmd/plan_server.go;
grpc/api/proto/forge/v1/forge.proto (ProposePlan, GetPlan, ProposedStep);
flutter/lib/features/forge/project/{project_loop,proposed_steps}.dart.