Skip to content

Judgements

A judgement is CORE’s independent verdict on a finding that an external assessment platform submitted. The judge agent (grpc/agents/judge) makes it and lands it on the finding in the console’s store. The page is DataEx › Judgements (?tab=judgements).

As deployed, the docket cannot be populated. CORE_JUDGEMENT_INGEST_TOKEN gates POST /api/judgement/submissions, and while it is unset the endpoint refuses everything. No manifest and no workflow in this repo provides it, and nothing in FACE or PULSE posts there. So the judge reports unable-to-judge indefinitely. That is the design working, not a bug.

One document, three routes

The console owns the shape: judgementStore in operators/core/internal/console/judgement.go has version, revision, updatedAt, submitters{} and judgeRuns[]. Its source of truth is the console’s judgement-submissions table. The core-system/core-judgement-submissions ConfigMap is a verbatim projection, rewritten after every commit, so never edit the ConfigMap by hand. The agent is a reader of that projection and defines nothing.

The routes are split by authority, never by a field in the body:

RouteMethodAuthorityWhat it does
/api/judgement[?submitter=NAME][&batch=ID]GETbrowser sessionthe submitted findings and CORE’s verdicts on them. Answers 200 even when the store is unreadable, and says so in the body
/api/judgement/submissionsPOSTthe submitter’s shared CORE_JUDGEMENT_INGEST_TOKEN as a bearerfindings in. Any verdict on the wire is discarded, because a platform may not grade its own work. The response is the stored, judged batch. Re-posting the same batchId is idempotent. The body limit is 1 MiB
/api/judgementsPOSTthe GitHub App’s rights for the console’s org (authorizeReporter)verdicts in, from CORE’s own judge agent. This is the only way a verdict enters the store
/api/judgementsGETbrowser sessionthe judge agent’s runs

The ingest token is refused by name at the judge’s door: “the submission ingest token has no authority here: a platform that submits findings does not judge them.”

Outcomes

concur, dissent or unable-to-judge (the console also accepts its own out-of-scope). Anything else is refused at the door. Unable-to-judge is never rendered as agreement. Evidence that is absent, unreadable, off-subject, stale or merely restates the claim, an unreachable model, an unparseable reply, and a submitter who said “undetermined” all produce an abstention with the reason.

The ladder is shared across platforms. It lives in github.com/org-runink/inference/judgement, and CORE builds it from .github/inference-snapshot. It runs eight deterministic gates before any model call. A quantitative claim such as null_rate <= 0.02 against a measured rate is settled by arithmetic, and the ladder recomputes the ratio from the evidence’s own numerator and denominator, dissenting when they disagree. The model sees only a qualitative claim, is asked one question, and is never told the submitter’s stance. The two are joined afterwards in Go.

Reason codes are stable strings, for example no-evidence, evidence-unreadable, evidence-about-another-subject, evidence-restates-the-claim, evidence-predates-the-staleness-horizon, evidence-has-a-zero-denominator, evidence-disagrees-with-itself, evidence-does-not-settle-the-claim, model-unavailable, model-response-unparseable and no-inference-endpoint-configured.

The fast path

Before its LLM call, the judge runs a small pinned tree model (judgefast, from ml/decide), switched by the Actions variable CORE_FAST_JUDGEMENT (default on). It answers only at the ladder’s model gate, never without evidence, and never on a health-adjacent claim. Whatever it does not answer goes to the LLM exactly as before. v0 abstains on almost every real claim, so verdicts today are still the LLM’s. Every model-based verdict records which path it took, for example basis: model (fast-model judgement-fast v0 sha256:…) or model (llm).

When the judge runs

.github/workflows/core-agent-judge.yml:

  • a schedule, 29 8 * * * (daily 08:29 UTC);
  • an @core_judgement comment;
  • workflow_dispatch;
  • repository_dispatch: core-playbook.

It writes to two sinks: the console store (POST /api/judgements) and the living “⚖️ Judgement status” issue. A successful store write raises the harness event judgements.reported (dissent, unable).

Not the same as

  • agents/eval’s judge (eval/judge.go) scores our own agents’ run traces 0..1 against a golden set. agents/judge assesses somebody else’s finding against its evidence.
  • The plan gate in fixer, selfheal and triage runs the same ladder over each agent’s own proposal before it acts.