Skip to content

The planner

The model proposes; you approve. A free-form chat message goes to ForgeService.ProposePlan. The server hands it to the platform’s own in-cluster model and returns a ticket at once. The studio then polls GetPlan.

How a plan is made

  • The model is sovereign. The planner sends a plain net/http request to INFERENCE_URL/v1/chat/completions with the model name from INFERENCE_MODEL. There is no SDK and no external endpoint. INFERENCE_URL should point at the general tier; the coder tier is for the coding agents.
  • It is slow on purpose. Planning runs at CPU decode speed. Each plan runs under its own 12-minute context with max_tokens 500.
  • One plan at a time. A second ProposePlan while one is thinking fails immediately with ResourceExhausted: the planner is still thinking about the previous message — the platform's own model decodes one plan at a time on CPU; wait for that plan, then ask again. It is not queued.
  • No grammar-constrained decoding. The model is asked for compact JSON. The object is extracted (prose and code fences are tolerated) and then validated.
  • Plans live in memory only. They are kept for one hour after finishing. A restart forgets them, and GetPlan then answers NotFound: the planner forgot this plan (restart or expiry) — ask again. Nothing about a plan goes to GitHub or appfs, and ProposePlan writes no audit entry.

A message may be up to 4000 characters. An empty one gets say what you want the forge to do.

The answer is untrusted

The server re-validates every step. Invalid steps are dropped, never repaired, and the reply says how many were dropped:

  • the action is new or change;
  • the name follows the repo-name rule; a change must name an existing app, and a new must not;
  • the kind maps to an AppKind;
  • the brief is within the same length rules as a filed brief;
  • there are at most 5 steps.

Errors from the inference plane reach Plan.error unparaphrased.

Every step is judged before you see it

FORGE mounts the shared Runink plan gate (inference/judgement JudgePlan). The verdict goes on ProposedStep.judgement (runink.ui.judgement.v1.Judgement). The evidence is your message, the organisation’s forged apps (when GitHub could be read) and the selected app. The claim is always “the owner asked for this step”, never the model’s own rationale.

VerdictWhat the studio shows
ConcurThe step, as a normal proposal
DissentMarked “judge dissents” with the reason, accepted = false, offered only as Use anyway
Unable to judgeMarked as such. Never shown as agreement
UnsetNot judged (FORGE_JUDGEMENT=off)

Any verdict the model wrote itself is stripped before judging, and the reply says how many. The brief itself is never altered: it reaches the agent verbatim.

Approval

Whether it was judged or not, a proposed step is only a prefill. Choosing it puts a typed command in the console input, for example new pipeline <name>: <brief> or change <name>: <brief>. You may edit it, and nothing happens until you send it. The send goes through the same deterministic command path as anything you type, and the audit records it as yours.

Sources

grpc/internal/planner/{planner,model,judge}.go; grpc/cmd/plan_server.go; grpc/api/proto/forge/v1/forge.proto (ProposePlan, GetPlan, ProposedStep); flutter/lib/features/forge/project/{project_loop,proposed_steps}.dart.