Skip to content

Playbooks

Intelligence › Orchestrator (?tab=orchestrator) is where you manage playbooks. A playbook is a chain of steps with the stage labels detect → validate → remediate → notify → verify. It links a playbook to CORE’s own agents in both directions:

  • a playbook step can start an agent;
  • an agent’s report can start a playbook.

Atlas listed hard-coded playbooks with typed-in success rates and buttons that did nothing. CORE makes each of those real or says it is absent.

  • Contract: PlaybookService in grpc/operators/core/api/proto/runink/core/atlas/v1/playbook.proto
  • Service: internal/console/atlas_pb.go
  • Store: atlas_pb_store.go
  • Executor and ticker: atlas_pb_engine.go
  • Agent doors: atlas_pb_http.go
  • Cron: atlas_pb_cron.go

Lifecycle

PlaybookStatusHow it gets there
DRAFTEvery playbook is created as a draft. Editing an ACTIVE playbook returns it to DRAFT
ACTIVEOnly through ActivatePlaybook, by an Atlas admin who neither wrote nor last edited it (four eyes). An active playbook starts agents on its own
PAUSEDPausePlaybook. ResumePlaybook brings it back

Clients may not set status: status is server-managed: a playbook is written as DRAFT and becomes ACTIVE only by ActivatePlaybook (four eyes).

DraftFromFindings builds a DRAFT deterministically from the current CapEx findings (the rules failing most) as capex_scan → approval → notify → capex_scan (verify). No model is involved.

Caps: 100 playbooks, 1..10 steps, 4 triggers per playbook.

Steps

A step has a kind (the Atlas stage label: DETECT, VALIDATE, REMEDIATE, NOTIFY or VERIFY) and exactly one action. The kind is a label for people. The action is what runs.

ActionWhat it doesBounds
dispatch_agentSends repository_dispatch event_type core-playbook to CORE’s agents repository, then waits for a step report until the deadline. No report by then ends the step TIMED_OUT. A step is never assumed successfulagent must be dispatchable. timeout_minutes 5..240, default 60
capex_scanRe-evaluates the current CapEx feeds and records a ScanRun (trigger MANUAL). With compare_to_first_scan, it passes only if the scored finding count (optionally for rule_ids) did not rise compared with the run’s first scanNeeds an earlier capex_scan to compare to
require_approvalPauses the run until DecideApproval. An empty approver means any Atlas admin, and it is never the run’s initiator. Expiry ends the run EXPIRED. Rejection ends it REJECTEDtimeout_hours 1..168, default 72
notifyPosts one comment of ids and counts on a GitHub issue in the console’s org. It never includes feed rows or finding details. Nothing else is ever contactedrepo is a repository name (no owner, no URL). issue is a positive number

What no step does: call a model, edit a feed or source system, or contact a registered enterprise agent. expected_lift is always absent, because it would be a forecast CORE cannot make. success_rate is measured as SUCCEEDED ÷ finished runs, and is absent until a run has finished.

Steps are claimed by CAS before they run and are never re-run. An abandoned claim ends FAILED.

Triggers

TriggerFires when
manualOnly RunNow starts it
scheduleA 5-field cron in UTC whose minute field is ONE number from 0 to 59, so it fires at most once an hour. Anything else is refused: the minute field must be ONE number 0-59 — a schedule fires at most once an hour
agent_eventAn agent’s event matches every condition (AND). An empty condition list matches every event of that agent and event
capex_eventThe CapEx engine raises scan.committed, the only CapEx event

cooldown_minutes sets the minimum time between runs started by the same trigger: 5 to 43,200 (30 days), default 60 (atlas_pb.go). A condition is {attribute, op, value}, where op is one of eq, neq, gt, gte or contains. A numeric op on a non-number evaluates to false, never to an error.

Events the harness raises

Events fire only after a successful store write, never before. No agent change is needed for any of these.

AgentEventRaised byAttributes
datagovfindings.reportedPOST /api/data-governance-findingsfail_count, warn_count, and per fail/warn finding control, verdict, domain, connection
reconrecon.reportedPOST /api/rules-recon/reportdrift, shadow, missing, tenant
judgejudgements.reportedPOST /api/judgementsdissent, unable
capexscan.committeda committed CapEx ingestcritical, high, medium, findings, dq, rule_id (matches when that rule has at least one scored finding)
estateestate.mappeda committed Resolve estate mapdomains, unplaced, gaps, drift, shadow, missing

Any agent may also raise an explicit event through POST /api/playbook-events (below).

The loop guard

An agent that a playbook dispatched will, by doing its job, post the very report that could re-trigger the playbook. So an event is skipped, and the skip is recorded on the event with its reason, when any of these holds:

  • the playbook is not ACTIVE;
  • a run is already in flight (a run is in flight (…));
  • the trigger’s cooldown has not passed (cooldown: trigger … started a run at …);
  • the chain depth exceeds 3 (chain depth … exceeds 3 (caused by run …) — refused, not run).

ListEvents shows what each event started or skipped, and why.

Dispatching agents

Only six agents can be named in dispatch_agent. These are the autonomous ones that need no PR or issue context (pbDispatchable):

datagov · compliance · risk · judge · recon · curator

ListDispatchableAgents returns that set, plus whether the console can dispatch at all right now (a GitHub App is configured and an agents repository is resolved) and why not.

The dispatch goes to <App installation org>/<CORE_AGENTS_REPO>, where CORE_AGENTS_REPO defaults to core. It uses the console’s App installation token. It is a repository_dispatch and not a workflow_dispatch, because the App holds actions: read only.

{
  "event_type": "core-playbook",
  "client_payload": {
    "agent": "datagov",
    "playbook_id": "…",
    "run_id": "…",
    "step_id": "…",
    "chain_depth": 0
  }
}

Each of the six workflows accepts repository_dispatch: [core-playbook] and runs only when client_payload.agent names its own agent. .github/scripts/playbook-context.sh validates the payload before the agent sees it: ids must match ^[A-Za-z0-9._-]{1,64}$ and depth must be 0..3.

The job’s last step, .github/scripts/playbook-step-report.sh, reports the outcome back. The reported conclusion is the agent step’s outcome, which the job status can make worse but never better. The report is retried on 5xx and network errors. A 4xx is logged and not retried. An undelivered report fails the workflow run, and the console ends that step TIMED_OUT.

A curator dispatched by a playbook is treated exactly like its scheduled run (fleet-wide, with the same dry-run rule), so a playbook can never make it publish more than its cron would. See the agent fleet.

The two agent doors (JSON)

Both doors take the same bearer contract as the agents’ other report ingests: a GitHub App token that carries the App’s authority for the console’s org, verified by authorizeReporter. Bodies are strict, text is screened, and every call is audited.

POST /api/playbook-runs/step-report

{"runId": "…", "stepId": "…", "agent": "datagov",
 "conclusion": "success|failure|cancelled|skipped",
 "detail": "one line", "githubRunUrl": "https://github.com/…/runs/N"}
StatusMeaning
200{"accepted": true}
400Malformed body
401missing bearer token
403The token is not the App’s authority for this org
404Unknown run or step
409The step is not awaiting a report, or it is another agent’s step

The console, not the agent, decides what the conclusion does to the run.

POST /api/playbook-events

{"agent": "datagov", "event": "findings.reported", "subject": "…",
 "detail": "…", "attributes": {"control": "DG-3", "verdict": "fail"},
 "causedByRunId": "…"}

It answers 200 {"started": ["PBR-…"], "skipped": [{"playbookId": "…", "reason": "…"}]}. Attributes are string→string, capped at 16 keys of 256 bytes each, and screened like every other free text. They are matched against trigger conditions and never executed. causedByRunId sets the chain depth.

Runs

RunStatus:

StatusMeaning
RUNNINGA step is executing
WAITING_AGENTWaiting for a dispatched agent’s step report
WAITING_APPROVALWaiting at a require_approval step
SUCCEEDEDEvery step finished successfully
FAILEDA step failed
REJECTEDAn approval was rejected
CANCELLEDCancelRun was called (note required)
TIMED_OUTA dispatched agent did not report by the deadline
EXPIREDAn approval was not decided in time

Triggers are MANUAL, SCHEDULE, AGENT_EVENT and CAPEX_EVENT.

The ticker (Start, about every 30 s with jitter) fires cron slots, deadlines and expiries. Each is claimed by CAS, so replicas sharing one appfs root act once.

The whole state is one metadoc, var/lib/atlas-playbooks, with one CAS per state transition. It holds:

  • 100 playbooks;
  • every in-flight run plus the newest 200 finished runs;
  • 300 events.

It has a 768 KiB budget. When the budget is reached it first trims history and then refuses.

RunNow starts a run of an ACTIVE playbook by hand. It fails with FAILED_PRECONDITION when the playbook is not ACTIVE or a run is in flight.

Who may do what

RPCClassAudit action
ListPlaybooks, GetPlaybook, ListRuns, GetRun, ListDispatchableAgents, ListEventsread—
CreatePlaybook, UpdatePlaybook, DeletePlaybook, DraftFromFindings, PausePlaybook, ResumePlaybook, RunNow, CancelRunadmin-writeatlas.playbooks.create / .update / .delete / .draft / .pause / .resume / .run / .run.cancel
ActivatePlaybookdecider-write: an Atlas admin who is not the author or last editoratlas.playbooks.activate
DecideApprovaldecider-write: the named approver or an Atlas admin, never the run’s initiatoratlas.playbooks.approval.decide

Decisions, dispatches and reports also land in the governance decision log (subject playbook).