Skip to content

Agent Fleet

CORE runs its own agents on its own model. Two numbers describe the fleet, and both are correct:

  • Thirteen agents/ binaries under grpc/agents/: compliance, curator, datagov, deployer, eval, fixer, judge, opsdoctor, recon, reviewer, risk, selfheal, triage. Each is a stdlib-only Go module (its own go.mod, no SDKs), built from source per run.
  • Seventeen entries on the console roster (GET /api/agents, operators/core/internal/console/agents.go). These are the thirteen, plus four with no agents/ directory: resolver (a workflow around core session), forger (core-forge-run.yml), users (core-agent-users.yml) and session (the core session CLI). TestRosterCoversTheShippedFleet fails if one is dropped.

Before you build a new agent, read these. Several have nearly been reimplemented by people who did not know they existed. Each is one main.go you can read in one sitting.

The workflow file is the authoritative trigger, not the agent’s own header comment. Some headers still say “runs as a Cloud Run Job”, which has been false since Cloud Run was retired.

The roster

Roster nameDeliveryTrigger (from the workflow)What it produces
reviewercore-code-review.ymlpull_request opened/synchronize; a PR comment containing @core_review, @core_check, @code_review or @review; repository_dispatch: core-review from the org-wide sweepa ✔️ ack comment and core-review label, then PR review comments; findings set review:findings-open
fixercore-agent-fixer.yml (and inline in the review)runs inline on every finding; a @core_fix PR comment; workflow_dispatch -f pr=a draft correction PR onto the reviewed branch. Never merges
resolvercore-agent-resolver.ymlrepository_dispatch: core-resolve; cron 17 */3 * * *; workflow_dispatcha core session on the PR head, then a draft PR onto it
triagecore-agent-triage.ymlissues: opened; workflow_dispatcha triage comment plus labels
self-heal (agents/selfheal)core-agent-selfheal.ymlworkflow_run: completed on eight named workflows; workflow_dispatch -f run_id=one root-cause comment on the PR. Never pushes code
riskcore-agent-risk.ymlcron 23 5 * * *; workflow_dispatch; repository_dispatch: core-playbookthe living “🛡️ Dependency risk report” issue from core scan --json
compliancecore-agent-compliance.ymlcron 13 1 * * *; workflow_dispatch; repository_dispatch: core-playbookthe living “📋 Compliance status” issue, a control-evidence artifact, and a customer-safe block for release notes
curatorcore-curate.ymlcron 17 6 * * *; a core @curate / @curate comment; workflow_dispatch; repository_dispatch: core-playbookrelease notes (docs/release-notes.md), doc fixes, audits and test scenarios as a curator/<sha>-<runid> PR
deployercore-agent-deployer.ymlan @core_deploy issue comment; workflow_dispatch -f issue=the finished deliverable posted back on the issue
datagovcore-agent-datagov.ymlcron 11 7 * * *; @core_datagov; workflow_dispatch; repository_dispatch: core-playbookthe living “🗃️ Data governance status” issue and POST /api/data-governance-findings. No model
judgecore-agent-judge.ymlcron 29 8 * * *; @core_judgement; workflow_dispatch; repository_dispatch: core-playbookthe living “⚖️ Judgement status” issue and POST /api/judgements
reconcore-agent-recon.ymlcron 37 9 * * *; @core_recon; workflow_dispatch; repository_dispatch: core-playbookPOST /api/rules-recon/report, and the “🧭 Rules reconciliation status” issue when GITHUB_TOKEN/GH_REPO are set
opsdoctork8s CronJobevery 30 minutes (infrastructure/apps/github-runners/kustomize/opsdoctor/cronjob.yaml)comments on one tracking issue; its one repair is deleting a wedged runner pod
evalCLI, plus core-inference-bench.ymlmanual: score, replay, regress, model, bencha trajectory + outcome report and a gate exit code; bench prints TTFT and decode tok/s
userscore-agent-users.ymlan issue labelled add-users (the core @add_users form), opened by an org memberpatches spec.users on a ClientInstance, or rebuilds the demo allowlist
forgercore-forge-run.ymlcron 41 3,9,15,21 * * *; workflow_dispatcha core session against the oldest open forge brief, as a draft PR. Disarmed by default
sessionCLIcore session startsee Coding Sessions

The roster also carries a family (delivery or governance) and a delivery kind (workflow, cli, cronjob). The governance screens take their population from the family. risk, compliance, datagov, judge and recon are governance.

Models and budgets

All thirteen binaries default their model field to the literal "default". The reviewer and fixer workflows set REVIEW_MODEL=coder. Every model-calling harness agent goes through cmd/modelrouter, which is the one admission queue in front of the engines, on the background lane. Two stay direct to the engine on purpose: opsdoctor (it must not depend on the hop it reports on) and eval bench (it measures the engine itself).

Agent cost is decode-seconds, not runner-minutes. The model plane serves one request at a time. Every budget derives from one configured rate, INFERENCE_DECODE_TOK_S. Each agent’s budget.go defaults it to 3.4 tok/s, and the workflows pass vars.INFERENCE_DECODE_TOK_S. answerBudget(windowSeconds) derives max_tokens from the deadline rather than declaring it separately. At that rate a 2,000-token answer is about ten minutes of pure decode.

Size budgets from the deployment target, not the dev workstation. The 3.4 tok/s default is an estimate for the target box that has not yet been measured.

Switches

Variable (Actions variable of the same name)DefaultEffect
CORE_FIXER_JUDGEMENT, CORE_SELFHEAL_JUDGEMENT, CORE_TRIAGE_JUDGEMENTonthe shared plan gate (inference/judgement.JudgePlan) judges the agent’s own proposal before it acts. Only off, false, 0, no or disabled turns one off
CORE_FAST_JUDGEMENTonthe judgement fast path (judgefast) runs in front of the LLM for judge and the three plan gates
CURATOR_CRON_DRY_RUNunset means drythe scheduled curator publishes only when this is the literal string false
CORE_FORGE_ENABLEDoffarms the forger

With the plan gate, a dissent blocks the action. The fixer pushes nothing and reports unresolved, self-heal posts a “withheld” notice, and triage posts the classification with labels withheld. Unable-to-judge does not block: the fixer’s draft is titled [not verified by the judge], and self-heal and triage say NOT VERIFIED. Every run prints a line beginning PLAN-GATE: (off, empty, concur, dissent or unable) with the tally.

A green curator run proves nothing. A dry run does everything except publish and then finishes green, and an unset CURATOR_CRON_DRY_RUN means dry.

Started by Atlas playbooks

Six agents accept repository_dispatch event type core-playbook: datagov, compliance, risk, judge, recon and curator. Each job runs only when client_payload.agent names its own agent. The payload is validated by .github/scripts/playbook-context.sh and reaches the agent as PLAYBOOK_ID, PLAYBOOK_RUN_ID, PLAYBOOK_STEP_ID and PLAYBOOK_CHAIN_DEPTH. The job’s last step reports the outcome to POST /api/playbook-runs/step-report. A curator started by a playbook is treated exactly like the scheduled run. See Playbooks.

Pairs that look like duplicates and are not

PairThe boundary
curator security audit vs riskThe curator reads first-party code committed since its last run (CWE/OWASP, model-judged). risk runs govulncheck over dependencies (known CVEs, reachable).
curator compliance audit vs complianceThe curator starts from a diff and asks which controls it touches. compliance starts from the control index in docs/COMPLIANCE.md, checks every citation still exists, and deep-reads three controls.
selfheal vs opsdoctorselfheal needs a job that ran and failed. opsdoctor sees jobs that never run and runner pods that crash, from a CronJob outside Actions.
agents/selfheal vs mesh/selfhealThe agent diagnoses failing CI runs. mesh/selfheal is the circuit-breaker/remediator library the apps link.
agents/judge vs agents/eval’s judgeeval/judge.go scores our agents’ run traces against a golden set. agents/judge assesses someone else’s finding against its evidence.

Behaviours worth knowing

  • opsdoctor must never become a workflow. It heals the CI plane, so it cannot depend on it. It is the only agent that ships as an image (agents/ has exactly one Dockerfile).
  • compliance proposes and cannot apply. Its workflow holds contents: read.
  • datagov holds no credential and dials nothing. It assesses data quality and PII exposure only from what Resolve’s Explore recorded, read through GET /api/data-governance-estate. A source nobody explored stays unassessable. See Data governance.
  • judge abstains rather than agreeing. Its outcomes are concur, dissent and unable-to-judge. As deployed its docket stays empty unless CORE_JUDGEMENT_INGEST_TOKEN is set, because POST /api/judgement/submissions refuses everything while it is unset. See Judgements.
  • recon calls an ungoverned rule Shadow, so a fresh workspace is mostly Shadow. That is by design.
  • A model failure is loud. reviewer, triage and selfheal post a degraded notice when the model is unreachable instead of exiting green.

Managing schedules from the console

GET /api/agent-schedules lists fleet cadence and enablement, and which of it is editable. PUT /api/agent-schedules/{name} arms or disarms an agent with a body carrying at least one of enabled, dryRun or schedule. It is a privileged write gated by CONSOLE_AGENT_ADMINS and audited. In the console this is Agent runs › Runs, “Schedules & arming” tab (?tab=agentSchedules).