Skip to content
Independent Assessor

Independent Assessor

Roster name: judge. See also Judgements.

What it does

Before you act on a finding from another assessment platform, it reads the evidence itself and says whether it agrees, disagrees, or could not judge.

  • Judges findings that an external platform submitted to Runink TIDE.
  • Runs the same checks as a plan check on the Fix Drafter’s patches, CI Doctor’s diagnoses and Issue Triager’s labels before any of them acts.

What it reads

For each finding: the claim, and the evidence it relied on, such as excerpts, records and measured values. The model is never told the submitter’s own stance.

What it produces

One verdict per finding, with a reason: concur, dissent or unable to judge. Separately, the console marks a finding about a subject TIDE has no standing over as out of scope. Verdicts appear on the DataEx › Judgements page and in the living “Judgement status” issue. Each model-based verdict records which path produced it.

Human oversight

A person decides what to do with each verdict. A verdict does not change the finding’s owner system, and TIDE does not call the external platform back.

Model

  • Qwen3.6-35B-A3B, by Qwen, licensed Apache-2.0 (the general tier), for the qualitative remainder.
  • Qwen3-Reranker-0.6B, by Qwen, licensed Apache-2.0, to narrow long evidence before the language model sees it.
  • A small decision-tree model answers only the clear cases at the model step; anything it does not answer goes to the language model.

Runink domain adaptation for this agent is planned; this release uses the base model.

Where it runs and data handling

On your TIDE deployment’s own inference, on your Server or in your cloud. Findings and evidence go to that model plane and to no third-party AI service. Findings arrive through TIDE’s own intake.

Guardrails

  • Arithmetic before language: a measured claim is recomputed from the evidence’s own numbers, and a mismatch is a dissent.
  • Deterministic checks run before any model call.
  • Never agreement by default: missing, unreadable, off-subject or stale evidence, or an unreachable model, produces unable to judge with the reason.
  • Verdicts are never averaged into a score or an agreement rate.
  • Evidence is marked as untrusted data, with chat control sequences neutralised, and screened by the guardrails; flagged or unassessable evidence is unable to judge, never a verdict.

Limitations

  • It judges only the evidence submitted. It does not fetch more.
  • Subjects TIDE has no standing over are marked out of scope by the console, not judged.
  • The docket stays empty until a submitting platform is configured; see Judgements.

Evaluation

No published evaluation scores yet.

Illustrative example

Invented figures. A platform claims “the null rate of customer_id is at most 2%” with evidence of 310 nulls in 10,000 rows. TIDE recomputes 3.1% from those two counts and answers dissent: “The evidence shows a null rate of 3.1%, above the 2% claimed.” A second claim arrives with no evidence: unable to judge, “no evidence was submitted.”