Skip to content

CI Doctor

Roster names: self-heal and opsdoctor. See also the Agent Fleet.

What it does

When a build fails, it posts one root-cause comment with a proposed fix. When the build machines themselves stop working, it reports why and makes one bounded repair.

  • Self-heal: diagnoses a CI run that ran and failed.
  • Opsdoctor: checks the self-hosted build plane every 30 minutes for runners that crash, runners registered with no machine behind them, and jobs waiting for a runner that does not exist.

What it reads

  • Self-heal: the names of the failing jobs and the tail of each job’s log.
  • Opsdoctor: the state of the runner pods and the CI job queue, and the last log lines of a crashing pod in the namespaces it watches (the build plane and the platform’s own services). It reads no CI job logs.

What it produces

  • Self-heal: one comment that cites the job and the log lines, names the root cause and gives a short proposed diff. It goes on the pull request when the run has one, and otherwise on a tracking issue for that workflow. If the independent check dissents, a “withheld” notice posts instead. If it could not decide, the comment says NOT VERIFIED.
  • Opsdoctor: one comment on a tracking issue with the most likely cause and the next action, and the crash evidence it read, including log tails.
  • A “did not diagnose” notice, rather than a silent pass, when the model cannot be reached or returns nothing usable.

Human oversight

A person applies any code fix. Self-heal never pushes code. Opsdoctor’s only repair is deleting a crash-looping runner pod, at most two per run. Anything bigger goes to a person.

Model

  • Self-heal: Qwen3-Coder-30B-A3B, by Qwen, licensed Apache-2.0 (the coder tier).
  • Opsdoctor: Qwen3.6-35B-A3B, by Qwen, licensed Apache-2.0 (the general tier).

Runink domain adaptation for this agent is planned; this release uses the base model.

Where it runs and data handling

On your Runink TIDE deployment’s own inference, on your Server or in your cloud. Build logs go to that model plane and to no third-party AI service. Opsdoctor runs outside CI on purpose, so it can still report when CI is down, and it calls the inference engine directly.

Guardrails

  • One comment per failure, checked first: a wrong root cause sends the author the wrong way, so a dissent withholds it.
  • Logs are evidence, never instructions. A failing test can print anything, so log text is marked as untrusted data, with chat control sequences neutralised. Self-heal also screens the logs and its answer with the guardrails; opsdoctor does not.
  • Opsdoctor is asked to name the log or state it would need when the evidence is thin, instead of guessing.
  • Bounded repair: at most two runner-pod deletions per run, only for pods the cluster reports as crash-looping, only in the runner namespace.

Limitations

  • Self-heal needs a job that ran and failed. Jobs that never started are opsdoctor’s.
  • It does not fix the code itself, and it does not report delivery metrics.
  • A failure caused by an account billing or quota stop is skipped, not diagnosed.

Evaluation

No published evaluation scores yet.

Illustrative example

Invented failure. A test fails with want "degraded", got "healthy". Self-heal comments: “Root cause: the new order checks readiness before restarts, so a crash-looping pod reads as healthy. Fix: move the restart check first (diff attached).” The author applies it.