--- name: doctrine-debug description: Use when the user reports something broken, throwing, failing, or slow and wants it diagnosed and fixed with the doctrine: parallel hypothesis agents, red-teaming the diagnosis, and looping until the signal stays green. --- # Doctrine Debug Debugging under the doctrine. **REQUIRED BACKGROUND:** Read the `doctrine` skill first. The core discipline is Matt Pocock's `diagnosing-bugs` skill: follow its phases exactly; this wrapper only adds the doctrine posture around it. The diagnosis phases (loop-building through hypothesis testing) are ONE doctrine phase; the fix is a second. The diagnosis phase exits when the root-cause claim survives both red-team refutation and your own source-level trace. **Doctrine step 5 lets a wrapper replace the exit condition *and its counters*, so here are this phase's:** count hypothesis waves, add 1 for every wave that did not end in a confirmed root cause — one where every hypothesis died on evidence, and equally one where a hypothesis survived without becoming provable or the evidence came back inconclusive. **Counting only the waves that killed everything is how this counter never fires**: the wave that eliminates nothing is the worse sign of the two, it is the one a stuck diagnosis repeats, and a phase made of them can run without bound while the fix phase's own round alarm never starts because there is no fix phase yet. Escalate to the user as doctrine step 5's round alarm does, with this count in place of its round count. Without it, that alarm never fires here: it counts review rounds that found a blocker, and this phase runs no gate at all, so the stop is absent from the phase most likely to grind. Doctrine step 5's time alarm applies to this phase as written. **Keep that count on disk** with whatever doctrine step 5 keeps there for the fix phase, since a compaction mid-diagnosis takes an in-context count and announces the loss to nobody. Doctrine step 5's gate applies to the fix phase as written. Its designated review skill is `matts-code-review` on the fix diff — this wrapper already scopes the diff, and two things finish the job. **Pin the fixed point yourself: the SHA the fix phase started from** (`git rev-parse HEAD` before you write the regression test). Dispatched without one it stops and asks the user, twice a phase; given `main` or the merge-base it reviews everything already merged and re-files those findings, which doctrine step 5 counts as outstanding against the pass you are in. And **give it a spec source** — the bug's tracker issue, or the root-cause claim you wrote down in step 4 — or its Spec sub-agent skips with "no spec available" and the gate's designated review runs one axis of two. (Issue references resolve only where Matt's `docs/agents/issue-tracker.md` exists; without it, pass the file path.) ## Flow 1. Ask the user any open questions (repro steps, when it last worked, environment). 2. Build the feedback loop per `diagnosing-bugs` Phase 1. This is the skill: spend disproportionate effort here. If no local repro seam exists (no local DB, auth-blocked UI), use the project's documented CI or prod-probe paths, or stop and ask; don't hypothesize without a loop. 2b. Where the loop runs on a device or a host you have to log into — a router, a switch, a console, a box over SSH — **and the user has asked for a pane**, invoke `doctrine-pane` and drive it there, so the user watches the session and can take the keyboard when your hypothesis is wrong in a way only they can see. The request is the trigger and there is no heuristic behind it: `doctrine-pane` forbids opening one nobody asked for, so a step that infers the need is a step that skill will refuse. Unasked, run the command non-interactively and say in the report that a pane would have shown more. It is the seam this wrapper reaches for most, because a bug you cannot reproduce locally is usually one that lives somewhere you have to log in to. Not installed, or no herdr: say so in the report, run the command non-interactively, and record the interactive step as **not run** (doctrine step 3) rather than as a step that passed. Only the orchestrator opens a pane, never a dispatched probe. 3. When multiple hypotheses are live, test them as a parallel wave: one agent per hypothesis, each reporting evidence for/against. **Assemble each probe's prompt (doctrine step 2) around the feedback loop step 2 built** — the exact command, the invocation and output you already have from it, the symptom it asserts, the single hypothesis that agent owns, **and its isolation mode — read-only stated as a constraint in the prompt itself**, because worktrees and separate instances you can impose from outside and read-only you can only instruct. `diagnosing-bugs` Phase 1 does not end until you can name **one** already-run, red-capable command, so the artifact exists; handing it over is the whole reason for having built it. A probe given "test hypothesis 3" and nothing to run reads code instead and returns code-reading dressed as evidence — which "kill hypotheses on evidence, not vibes" then judges, and step 4 red-teams a diagnosis built on it. Probes must be isolated (read-only analysis, separate harness instances, or worktrees) so one-variable-at-a-time still holds. Kill hypotheses on evidence, not vibes. If every hypothesis dies, regenerate from the new evidence; don't recycle the old list. 4. Red-team the surviving diagnosis BEFORE writing the regression test or fix: give the red team (doctrine step 4) the symptom, the feedback loop, and your root-cause claim; ask it to refute. Verify its counter-claims from source. No point locking in a test for a refuted diagnosis. **Write the surviving claim down where the fix lands** (doctrine step 1) — one paragraph: the symptom, the loop command, the root cause, the source evidence. It is the diagnosis phase's deliverable, it is what a fresh context after a compaction resumes from alongside the run-state file doctrine step 5 has you keeping from wave 1, and it is the spec source the fix phase hands `matts-code-review`. 5. Fix per `diagnosing-bugs` Phase 5 (regression test first, minimal fix), then a simplification review of the fix (doctrine step 6), finished and committed before the pass's native checks start, and of each repair before the pass that re-checks it: a bug fix that adds a new abstraction is usually the wrong fix. Then loop: feedback loop green + native checks + `matts-code-review` + red team on the diff, to doctrine step 5's exit condition for this phase. For an intermittent bug, "green" means the full stress run clean (sized against **the reproduction rate `diagnosing-bugs` Phase 1 pinned** — its completion criterion requires a deterministic loop or, for a flaky bug, a pinned high rate raised until the bug is debuggable, so that number exists and is the one to size iterations against), not one lucky pass and not a rate nobody measured. 6. Deliver per doctrine step 7. ## Red flags - Fix written before the feedback loop goes red on this bug. - Root cause accepted because the red team agreed, without a source-level trace. - A flaky bug declared fixed after a single green run. - A probe wave dispatched with "test hypothesis N" and no loop command; what comes back is code-reading with the word "evidence" on it. - Hypothesis waves that all died on evidence, uncounted, with the user never asked whether to keep going. - `matts-code-review` run against `main` on the fix diff, re-filing findings from work that shipped weeks ago.