Playbooks
Intelligence › Orchestrator (?tab=orchestrator) is where you manage playbooks. A playbook is a chain of steps with the stage labels detect → validate → remediate → notify → verify. It links a playbook to CORE’s own agents in both directions:
- a playbook step can start an agent;
- an agent’s report can start a playbook.
Atlas listed hard-coded playbooks with typed-in success rates and buttons that did nothing. CORE makes each of those real or says it is absent.
- Contract:
PlaybookServiceingrpc/operators/core/api/proto/runink/core/atlas/v1/playbook.proto - Service:
internal/console/atlas_pb.go - Store:
atlas_pb_store.go - Executor and ticker:
atlas_pb_engine.go - Agent doors:
atlas_pb_http.go - Cron:
atlas_pb_cron.go
Lifecycle
PlaybookStatus | How it gets there |
|---|---|
DRAFT | Every playbook is created as a draft. Editing an ACTIVE playbook returns it to DRAFT |
ACTIVE | Only through ActivatePlaybook, by an Atlas admin who neither wrote nor last edited it (four eyes). An active playbook starts agents on its own |
PAUSED | PausePlaybook. ResumePlaybook brings it back |
Clients may not set status: status is server-managed: a playbook is written as DRAFT and becomes ACTIVE only by ActivatePlaybook (four eyes).
DraftFromFindings builds a DRAFT deterministically from the current CapEx findings (the rules failing most) as capex_scan → approval → notify → capex_scan (verify). No model is involved.
Caps: 100 playbooks, 1..10 steps, 4 triggers per playbook.
Steps
A step has a kind (the Atlas stage label: DETECT, VALIDATE, REMEDIATE, NOTIFY or VERIFY) and exactly one action. The kind is a label for people. The action is what runs.
| Action | What it does | Bounds |
|---|---|---|
dispatch_agent | Sends repository_dispatch event_type core-playbook to CORE’s agents repository, then waits for a step report until the deadline. No report by then ends the step TIMED_OUT. A step is never assumed successful | agent must be dispatchable. timeout_minutes 5..240, default 60 |
capex_scan | Re-evaluates the current CapEx feeds and records a ScanRun (trigger MANUAL). With compare_to_first_scan, it passes only if the scored finding count (optionally for rule_ids) did not rise compared with the run’s first scan | Needs an earlier capex_scan to compare to |
require_approval | Pauses the run until DecideApproval. An empty approver means any Atlas admin, and it is never the run’s initiator. Expiry ends the run EXPIRED. Rejection ends it REJECTED | timeout_hours 1..168, default 72 |
notify | Posts one comment of ids and counts on a GitHub issue in the console’s org. It never includes feed rows or finding details. Nothing else is ever contacted | repo is a repository name (no owner, no URL). issue is a positive number |
What no step does: call a model, edit a feed or source system, or contact a registered enterprise agent. expected_lift is always absent, because it would be a forecast CORE cannot make. success_rate is measured as SUCCEEDED ÷ finished runs, and is absent until a run has finished.
Steps are claimed by CAS before they run and are never re-run. An abandoned claim ends FAILED.
Triggers
| Trigger | Fires when |
|---|---|
manual | Only RunNow starts it |
schedule | A 5-field cron in UTC whose minute field is ONE number from 0 to 59, so it fires at most once an hour. Anything else is refused: the minute field must be ONE number 0-59 — a schedule fires at most once an hour |
agent_event | An agent’s event matches every condition (AND). An empty condition list matches every event of that agent and event |
capex_event | The CapEx engine raises scan.committed, the only CapEx event |
cooldown_minutes sets the minimum time between runs started by the same trigger: 5 to 43,200 (30 days), default 60 (atlas_pb.go). A condition is {attribute, op, value}, where op is one of eq, neq, gt, gte or contains. A numeric op on a non-number evaluates to false, never to an error.
Events the harness raises
Events fire only after a successful store write, never before. No agent change is needed for any of these.
| Agent | Event | Raised by | Attributes |
|---|---|---|---|
datagov | findings.reported | POST /api/data-governance-findings | fail_count, warn_count, and per fail/warn finding control, verdict, domain, connection |
recon | recon.reported | POST /api/rules-recon/report | drift, shadow, missing, tenant |
judge | judgements.reported | POST /api/judgements | dissent, unable |
capex | scan.committed | a committed CapEx ingest | critical, high, medium, findings, dq, rule_id (matches when that rule has at least one scored finding) |
estate | estate.mapped | a committed Resolve estate map | domains, unplaced, gaps, drift, shadow, missing |
Any agent may also raise an explicit event through POST /api/playbook-events (below).
The loop guard
An agent that a playbook dispatched will, by doing its job, post the very report that could re-trigger the playbook. So an event is skipped, and the skip is recorded on the event with its reason, when any of these holds:
- the playbook is not
ACTIVE; - a run is already in flight (
a run is in flight (…)); - the trigger’s cooldown has not passed (
cooldown: trigger … started a run at …); - the chain depth exceeds 3 (
chain depth … exceeds 3 (caused by run …) — refused, not run).
ListEvents shows what each event started or skipped, and why.
Dispatching agents
Only six agents can be named in dispatch_agent. These are the autonomous ones that need no PR or issue context (pbDispatchable):
datagov · compliance · risk · judge · recon · curator
ListDispatchableAgents returns that set, plus whether the console can dispatch at all right now (a GitHub App is configured and an agents repository is resolved) and why not.
The dispatch goes to <App installation org>/<CORE_AGENTS_REPO>, where CORE_AGENTS_REPO defaults to core. It uses the console’s App installation token. It is a repository_dispatch and not a workflow_dispatch, because the App holds actions: read only.
{
"event_type": "core-playbook",
"client_payload": {
"agent": "datagov",
"playbook_id": "…",
"run_id": "…",
"step_id": "…",
"chain_depth": 0
}
}Each of the six workflows accepts repository_dispatch: [core-playbook] and runs only when client_payload.agent names its own agent. .github/scripts/playbook-context.sh validates the payload before the agent sees it: ids must match ^[A-Za-z0-9._-]{1,64}$ and depth must be 0..3.
The job’s last step, .github/scripts/playbook-step-report.sh, reports the outcome back. The reported conclusion is the agent step’s outcome, which the job status can make worse but never better. The report is retried on 5xx and network errors. A 4xx is logged and not retried. An undelivered report fails the workflow run, and the console ends that step TIMED_OUT.
A curator dispatched by a playbook is treated exactly like its scheduled run (fleet-wide, with the same dry-run rule), so a playbook can never make it publish more than its cron would. See the agent fleet.
The two agent doors (JSON)
Both doors take the same bearer contract as the agents’ other report ingests: a GitHub App token that carries the App’s authority for the console’s org, verified by authorizeReporter. Bodies are strict, text is screened, and every call is audited.
POST /api/playbook-runs/step-report
{"runId": "…", "stepId": "…", "agent": "datagov",
"conclusion": "success|failure|cancelled|skipped",
"detail": "one line", "githubRunUrl": "https://github.com/…/runs/N"}| Status | Meaning |
|---|---|
| 200 | {"accepted": true} |
| 400 | Malformed body |
| 401 | missing bearer token |
| 403 | The token is not the App’s authority for this org |
| 404 | Unknown run or step |
| 409 | The step is not awaiting a report, or it is another agent’s step |
The console, not the agent, decides what the conclusion does to the run.
POST /api/playbook-events
{"agent": "datagov", "event": "findings.reported", "subject": "…",
"detail": "…", "attributes": {"control": "DG-3", "verdict": "fail"},
"causedByRunId": "…"}It answers 200 {"started": ["PBR-…"], "skipped": [{"playbookId": "…", "reason": "…"}]}. Attributes are string→string, capped at 16 keys of 256 bytes each, and screened like every other free text. They are matched against trigger conditions and never executed. causedByRunId sets the chain depth.
Runs
RunStatus:
| Status | Meaning |
|---|---|
RUNNING | A step is executing |
WAITING_AGENT | Waiting for a dispatched agent’s step report |
WAITING_APPROVAL | Waiting at a require_approval step |
SUCCEEDED | Every step finished successfully |
FAILED | A step failed |
REJECTED | An approval was rejected |
CANCELLED | CancelRun was called (note required) |
TIMED_OUT | A dispatched agent did not report by the deadline |
EXPIRED | An approval was not decided in time |
Triggers are MANUAL, SCHEDULE, AGENT_EVENT and CAPEX_EVENT.
The ticker (Start, about every 30 s with jitter) fires cron slots, deadlines and expiries. Each is claimed by CAS, so replicas sharing one appfs root act once.
The whole state is one metadoc, var/lib/atlas-playbooks, with one CAS per state transition. It holds:
- 100 playbooks;
- every in-flight run plus the newest 200 finished runs;
- 300 events.
It has a 768 KiB budget. When the budget is reached it first trims history and then refuses.
RunNow starts a run of an ACTIVE playbook by hand. It fails with FAILED_PRECONDITION when the playbook is not ACTIVE or a run is in flight.
Who may do what
| RPC | Class | Audit action |
|---|---|---|
ListPlaybooks, GetPlaybook, ListRuns, GetRun, ListDispatchableAgents, ListEvents | read | — |
CreatePlaybook, UpdatePlaybook, DeletePlaybook, DraftFromFindings, PausePlaybook, ResumePlaybook, RunNow, CancelRun | admin-write | atlas.playbooks.create / .update / .delete / .draft / .pause / .resume / .run / .run.cancel |
ActivatePlaybook | decider-write: an Atlas admin who is not the author or last editor | atlas.playbooks.activate |
DecideApproval | decider-write: the named approver or an Atlas admin, never the run’s initiator | atlas.playbooks.approval.decide |
Decisions, dispatches and reports also land in the governance decision log (subject playbook).