--- name: supervise description: Dispatch a multi-wave plan to specialist agents with audit + risk-area abort. Supports --auto-push, --auto-merge, --goal-mode for budgeted runs, --verify-blocking for a hard completion-claim gate. NOT for writing the plan itself (that is /spec), and NOT for a single small edit with no waves — just make the edit. when_to_use: User has a written plan and says "run the plan", "/supervise ", or "full auto". tools: Bash, Read, Write, Edit, Grep, Glob --- # /supervise ## Goal Take a written plan (`~/.agent/plans/.md` with Wave 1..N sections) and run it end-to-end, dispatching the right specialist agent for each wave, auditing after each wave, and aborting on risk-area violations. ## Modes | Mode | Behavior | |---|---| | `/supervise ` | Default: full-auto. Dispatch, audit, advance. Stops on Wave fail or safeguard. | | `/supervise --goal-mode` | Tracks state in SQLite via `core/infra/supervisor-goal.sh`. Resumable across sessions. Requires `sqlite3` + `jq` (the script exits 127 without them — see README Prerequisites). | | `/supervise --auto-push` | Each wave commits + pushes + opens PR. User merges. | | `/supervise --auto-merge` | Each wave commits + pushes + admin-merges via `auto-ship.sh`. | | `/supervise --verify-blocking` | Each wave's 2d audit also runs `core/infra/completion-gate.sh` with `AGENT_VERIFY_BLOCKING=block` — a REFUTED completion claim STOPs the wave (default without this flag: `dryrun`, logs only, never blocks). | ## Permission friction (plan-scope-allow prerequisite) /supervise does not grant itself permissions. Whether wave edits hit the native permission prompt is decided by the `plan-scope-allow` gate (`docs/gate-registry.md`), which is active only when `AGENT_PLAN_ALLOW_MODE=on` is exported, or — with the env unset — when the workspace resolves to the `personal` trust tier (`docs/customization.md` § Trust tiers). In `collab` workspaces, expect a prompt per edit; prefer report-first waves there. Coverage is Write/Edit/MultiEdit only — Bash/MCP wave commands always keep their own gates and prompts (extension is backlog LE-1). Hard safeguards (risk-area abort, R4 mutex, gitleaks, test-failure abort) bind in every tier and every mode, including full-auto. ### Dispatch pre-flight (before any edit wave leaves the main loop) A dispatched subagent cannot answer a native permission prompt, and a **background** dispatch auto-denies any tool call that would prompt — so an edit wave dispatched without cleared permissions doesn't fail loudly, it silently loses its Edit/Write calls and reports garbage. Before dispatching an execution (file-editing) wave, confirm at least one of: 1. The plan-approval flag exists — `[[ -f /tmp/agent-plan-approved ]]` (written on ExitPlanMode approval; wiped at every SessionStart), which arms the `plan-scope-allow` gate above, **or** 2. The project's `.claude/settings.local.json` carries `Edit`/`Write`/ `MultiEdit` in `permissions.allow` (project-layer rules are the reliable layer — subagents do not dependably inherit user-level allows). If neither holds: dispatch the wave **foreground** (prompts then reach the user) or stop and tell the user which of the two to set up. Never send an edit wave to a background dispatch on an unverified permission surface. ## Model policy Who runs on which model — and what enforces it: | Work | Model | Enforced by | |---|---|---| | **Judgment** — planning/design, wave dispatch decisions (who does what), gate verdicts & abort/advance, result synthesis | The main session's top model. Runs in the main loop, or via an agent **without** a `model:` pin (inherit). Never dispatch judgment work to an agent pinned below the session model. | This rule — a convention (frontmatter absence = inherit; a call-time choice is not CI-checkable) | | Specialist dispatch (`code-reviewer` → sonnet, `security-reviewer` → opus) | Each agent's own `model:` frontmatter | Runtime applies frontmatter; `validate-plugin` CI drift guard keeps `agents/master-registry.json` in sync | | **Execution dispatch** — implementation waves | Workhorse (MID) tier, via an explicit `model` override on the Agent dispatch (no executor agent is shipped) | Delegation-contract `model` field (`skills/supervise/templates/delegation-contract.md`). CI guards the guardable half: the template's model field and reviewer/verifier read-only toolsets (registry-drift gate); the call-time override itself stays a convention | | Mechanical fixes (build/type/lint cleanup), lookups, fan-out workers | Low tier, via an explicit `model` override | Per-call override — a convention | The orchestrating session keeps judgment and dispatches hands: when a wave is execution work, dispatch it at the tier the table names instead of doing it inline at the session model. Inline execution at the top tier is the expensive default this rule exists to prevent. Two placement corollaries (economics: `docs/model-routing.md` → Intelligence placement & Floors): the audit-after-wave step is this harness's **advisor checkpoint** — TOP judgment re-ranking MID execution mid-run, which is where advisor value concentrates (not in a single upfront ranking). And a **coordination-cost check** applies before dispatching: a wave whose delegated volume would not clearly offset its own contract+report boundary (billed twice in each direction) is not a wave — fold it into an adjacent wave or make the edit directly. The supervise loop itself never overrides a model. `core/hooks/supervisor.py` is a dispatch-suggestion stub — it matches intent to a specialist from the registry; it does not read or set `model`. If you add an agent whose role is planning or deep design, leave `model:` out of its frontmatter. This table is the enforced (Claude) instance of the cross-runtime tier policy in `docs/model-routing.md` — see that document for the Codex/Gemini columns, the verify-judge floor, and the fan-out worker default. Prompt-side dispatch guidance for frontier models (anti-wrap-up, evidence-grounded claims, no reasoning replay) is advisory in `docs/concepts/fable-5-prompting.md`. ## Steps ### 0. Intake restatement Before touching the plan, restate the ask so dispatch decisions trace to a machine-checkable record rather than to chat prose: a. Fill `skills/supervise/templates/prompt-restatement.md` from the user's invocation text and the plan's objective — all six sections (Original ask verbatim / Interpreted goal / Assumptions / Out of scope / Success criteria / Open questions), plus the `Run started:` UTC timestamp line (the routing log accumulates across sessions; this is the audit window's start). b. Persist it to `.agent/plans//RESTATEMENT.md`. (RECORD.md stays a mechanical completion ledger — intake interpretation never goes there.) c. If **Open questions** is non-empty and the run is not full-auto, surface them to the user before Wave 1. Full-auto runs note the resolution chosen under Assumptions instead. `/manager-audit` grades this file (lane `restatement-quality`) — a wave serving a goal absent from **Interpreted goal** is flagged as scope drift. ### 1. Plan validation a. Read `~/.agent/plans/.md`. b. Confirm it has Wave 1..N sections. c. Run `bash core/infra/supervisor-goal-audit.sh score --plan --wave 1` — if verdict is `weak`, warn the user and ask whether to proceed. d. If `--goal-mode`, initialise: ```bash core/infra/supervisor-goal.sh init [] "" ``` ### 2. Per-wave loop For each wave i ∈ {1..N}: a. **Read Wave i section** of the plan. b. **Classify the wave and pick lanes** based on its content: - Judgment work (deciding, synthesizing) stays in the main loop. - Execution work dispatches with an explicit `model` per the Model policy (implementation → workhorse tier, mechanical cleanup → low tier), after the dispatch pre-flight above clears the permission surface. - Wave touches `core/hooks/` or general code → `code-reviewer` after - Wave touches auth/secrets → `security-reviewer` - High-stakes wave where a cross-vendor opinion is worth its cost → `/council-review` (skills/council-review/SKILL.md; paid, user-approved) - **Never route an execution wave to `code-reviewer` or `security-reviewer`** — both carry read-only toolsets (Read/Grep/Glob, CI-enforced); they cannot edit a file at all. `core/hooks/supervisor.py` suggestions name specialists for the *review* lane, not the execution lane. Every dispatch is written as a **delegation contract** — `skills/supervise/templates/delegation-contract.md` (goal / output format / tools & scope / boundaries, plus an explicit `model` field per the Model policy). Five orchestration rules travel with it (details in the template): - **Fan-out cap 3–5** per wave — a wave with more concurrent subtasks splits into consecutive waves (the template shows a worked split). - **Write single-threading** — one writer per fileset; review/verify agents carry read-only toolsets, which the CI registry-drift gate enforces. - **Self-contained contracts** — subagents inherit no conversation history; the contract carries every needed path, decision, and constraint, with the wave's relevant constraint slice re-stated (not whole rulebooks). - **Verifier isolation** — verifiers are fresh spawns with no author context or self-assessment; they grade end-state only. - **Worker reuse (cache)** — consecutive subtasks over the same fileset or context continue the *same* worker rather than fresh-spawning each one (a fresh spawn re-pays the full context write uncached). Verifier isolation is the standing exception: verifiers are always fresh. c. **Execute** the wave's intended changes — through the dispatched execution lane, not inline at the session model (inline is judgment's lane, not execution's). Under `--verify-blocking`, the wave's delegation contract (`skills/supervise/templates/delegation-contract.md` "Completion claim file") must instruct the worker to write `.agent/claims/-w.yml` (verify-completion schema) on finishing — completion-gate.sh in step 2d reads it, falling back to the repo-wide `.agent/claims/.yml` convention if the wave-suffixed path is absent. No instruction, no claim file: the gate still runs and still blocks in block mode, but on a "claim file missing" it cannot check anything against, not a real refutation — write the instruction so the gate has something to check. d. **Audit**: ```bash bash core/infra/supervisor-goal-audit.sh ``` - PASS: continue. - FAIL: STOP. Report to user. Do not auto-fix and retry — the user decides next action. Then, the **completion-claim gate** (`docs/gate-registry.md` `completion-gate` row — the first blocking consumer of `completion-verify.py`'s verdict, `docs/scoring-convention.md`): ```bash bash core/infra/completion-gate.sh [claim-path] ``` - `--verify-blocking` on the `/supervise` invocation sets `AGENT_VERIFY_BLOCKING=block` for this call (unset/other flags leave the script's own default, `dryrun` — logs the verdict, never blocks). - block mode + REFUTED (exit 1): STOP, same protocol as a 2d audit FAIL — report the printed refutations to the user verbatim, do not auto-fix and retry. The user decides next action. - exit 2 (the gate's own liveness canary failed — a dead completion-verify.py, not a refuted claim) is also a STOP: the gate cannot be trusted, report it as a harness defect, not a wave failure. e. **Advance** (if `--goal-mode`): ```bash core/infra/supervisor-goal.sh advance-wave ``` f. **Wrap** (if `--auto-push` or `--auto-merge`): - Invoke `/wrap` with the appropriate flag. ### 2b. Race lanes (opt-in: wave annotated `race: true`) For a high-stakes wave the plan may annotate `race: true`: the SAME spec goes to two independent implementation lanes, and the supervisor picks the stronger result. Rules that keep it inside the standing invariants: - **Two lanes, both patch-only — neither edits the tree.** Lane A is a Claude subagent in an isolated worktree; lane B is the `implementer` role via `core/infra/call-worker.sh` run from a dedicated `.worktrees/race-/` checkout (which also scopes its workspace-write sandbox). Each lane's deliverable is a patch file under `.agent/workers/`, produced with `git diff`, plus a lane report (`skills/supervise/templates/lane-report.md`). - **The supervisor is the only tree writer** (one-writer rule preserved): it runs the wave's acceptance command against each patch, compares, applies exactly ONE, and records which lane won and why in RECORD.md. - **Race = 2 slots against the wave's fan-out cap.** A race wave carries at most one other concurrent worker. - **Judgment stays home**: the compare-and-pick step is main-loop judgment work, never dispatched. - If a lane is unavailable (`status: unavailable` capture), the race degrades to a single-lane wave — stated in RECORD.md, never silently. ### 3. Safeguards (immediate abort) Stop the supervise loop immediately on any of these: 1. User says "stop" / "halt" / "cancel" / "pause". 2. Risk-area violation detected (`rules/policy/security-guards.md` 5 areas). 3. R4.1 file mutex blocked. 4. gitleaks failure at any stage. 5. Test suite failure. 6. Type-check / lint failure. 7. A dispatched review/verify agent died (session limit, API error) — treat the wave's audit as FAIL even if the aggregate reports 0 findings; a dead reviewer is not a clean review (re-dispatch or verify in the main loop). On safeguard: emit `blocked` broadcast (R13), report state to user. ### 4. Token budget (--goal-mode) If `--goal-mode` with a budget, after each tool call: ```bash core/infra/supervisor-goal.sh track-tokens ``` When `status` transitions to `budget_limited`, the helper auto-writes a graceful-wrap stub at `wiki/synthesis/-budget-limited-.md` and the supervise loop exits gracefully. ### 5. Completion When wave N passes audit: ```bash core/infra/supervisor-goal.sh complete # if --goal-mode ``` Report: ``` Plan : COMPLETE Waves: N/N PRs opened: PRs merged: Outstanding: ``` Also write the same facts to `.agent/plans//RECORD.md` — the **repo-native execution ledger** (waves / PRs / audit verdicts / carried items). The ledger is mechanical facts only; session narrative belongs to the global recording layer, so the two never duplicate. On `--goal-mode` runs the `complete` command drops a RECORD.md stub automatically (the deterministic guarantee the file exists — it never overwrites one you already wrote); on non-goal runs writing it is this step's discipline. This keeps an execution record on runtimes that have no global recording layer at all. Finally, auto-run the two model-economics lanes: `bash core/infra/manager-audit.sh --json --since `, then filter its `findings` array to `lane == "routing-waste"` or `lane == "token-spend"` (the script has no per-lane flag — the JSON `lane` field on each finding IS the selection interface; the other two lanes' findings are simply ignored in this pass). First check whether `.agent/logs/model-routing.jsonl` exists and has at least one record scoped to this run — if it's missing or empty (the `routing-log-missing` WARN finding, or a `routing-waste`/`token-spend` finding set with zero `RECORDS`), the correct summary is `routing: skip (no routing log)`, NOT `routing: clean` — a missing observer log makes the two lanes trivially FAIL/WARN-free (manager-audit.sh short-circuits to an empty `RECORDS` set and reports `lane-clean` PASS), so "clean" would misrepresent "never measured" as "measured and good." Only write `routing: clean` when records actually exist and both lanes are FAIL/WARN-free. Otherwise write `routing: N WARN/FAIL — top: [] rel_cost=` (from the `token-spend` lane's `top-spend-sources` finding). This is non-blocking — a bad verdict never blocks completion, it only informs the ledger. Then offer to run the FULL `/manager-audit ` (all four lanes — restatement quality, model-routing waste, relative token spend, role compliance), passing `--since ` so the audit scopes to this run's dispatches. It reads the logs this run already produced; it never blocks completion. ## Hard rules - **Never skip the audit step.** A wave isn't done until the audit passes. - **Never auto-retry a failed audit.** Hand off to the user. - **Never bypass safeguards** even if the user said "full auto". "Full auto" = no clarifying-Q at decision forks; it doesn't disable safety gates. - **Always emit broadcasts** at started / decision / committed / pr_opened / done / blocked (R12 / R13).