--- name: graph description: Work a plan held in the enforcer graph — claim the next runnable node yourself, do it, keep the lease alive, report it against its acceptance criteria with captured command output as evidence. Use when a project has .claude/graph.json, when asked to "work the plan", "take the next node", "claim from the frontier", when a coordinator hands you a graph node to work as a subagent, when fanning a plan out to subagents, or when the enforcer graph MCP tools (graph_next_work, graph_report, graph_heartbeat, graph_plan_status, graph_remember) are available. --- # Working a graph The enforcer graph holds a plan as a DAG. Nodes are tasks; each carries `data.acceptance`, the list of lines your report will be judged against. You never decide what to do next: the graph does, from the edges. ## The worker rule: whoever does the work claims and reports it The plugin's hooks capture evidence **per agent**. A subagent's commands go into that subagent's capture, and are attached only to a run **that same agent** claimed with `graph_next_work`. A run claimed by anyone else gets none of them. Measured: a coordinator claimed and reported 18 nodes on behalf of its subagents, relaying their summaries as evidence, and all 18 were judged `rejected / unsupported_by_evidence`. - **Claim your own node with `graph_next_work`**, from the agent that will run the commands. If you were told to work a particular node, treat that as a hint: when the card hands you a different node, a sibling already took yours. Work the node you hold and say which one it is. On MCP 0.9.5 or newer, a coordinator can pin the claim: `graph_next_work` with `node` (key or id) claims exactly that node when it is runnable, and answers `not_runnable` with the reason when it is not. That is still `graph_next_work`, so the hooks follow it. - **Never claim by id.** Do not claim through `enforcer_api_write` or `take_task`. The hooks only follow `graph_next_work`, so a run claimed any other way gets no heartbeat and no captured evidence. - **No `graph_next_work` in your tools?** Your token lacks the graph write scopes (or the server predates enforcer-v3-mcp 0.9.4, which stopped hiding the worker tools from fine-scoped tokens). Stop and say so; do not fall back to `enforcer_api_write`. `/enforcer:login` with the graph scopes fixes it. - **Report before you return.** The Stop guard does not run for subagents, so nothing reminds you. A run left open when you hand back is a lease that lapses with nothing to show for it. - **Never report a node you did not claim.** Never re-report or re-judge one. ## The loop 1. **`graph_next_work`** (`graph`, optional `for`, `runner`, `node`, `upstream_depth`) — claims one runnable node under a row lock and returns one card: the node, its `criteria` (its own acceptance lines, the graph and tenant mandates marked `[mandate: …]`, and `[check]` rows the server evaluates in code), its `inputs` (values upstream nodes produced), attached `skills` (read them before starting), recent `observations` (facts earlier runs left), and the `run` you now hold (`run_id`, `lease_expires_at`). Two sessions can never receive the same node. - **`upstream`** (MCP 0.9.5+): for each prerequisite, nearest first, the run's report, observations, links and file ids, one `line` each to read first. Read it before you start; do not re-derive what it already says. - **`route`** (graph with routing on): the agent profile the router recommends for this node, and why. Advisory. `route: null` means routing is on and nothing is recommended. See "Following the route" below. - `state: wait` — nothing runnable, but something is running, verifying, or time-gated. When the card has `wait_until`, nothing can start before that instant: say so and stop. Do not poll in a tight loop. - `state: complete` — nothing runnable, running, verifying or time-gated. The plan is done, or failed nodes block the rest; the card says which. An **open** graph never completes: it answers `idle` instead, which means "nothing to do right now", not "finished". - **Assignment:** when the project config sets `graph.assignment` (`me` or `mine`), pass it as `for`. `assigned_elsewhere` means only other people's nodes are runnable; assignment is advisory, so take one only if a human says to. 2. **Work the node** in its own worktree or branch named after `node.key`. 3. **`graph_heartbeat`** (`graph`, `node_id`, `run_id`) — you hold a lease, not the node. The plugin's hook heartbeats every ten tool calls for you. Call it yourself before any long silent step, such as a build, a long test run or a wait on CI. The response `state` is an instruction: - `ok` — keep working. - `cancel_requested` — stop, `graph_remember` what is worth keeping, then `graph_report` with `status: cancelled`. - `reclaimed` — your lease lapsed and another harness owns the node. STOP. A report from this run will be refused. `graph_remember` any progress worth keeping, then `graph_next_work`. - `finished` — the run already ended. Nothing to report; `graph_next_work`. A graph **reset** also ends every open run this way: stop, re-read the plan (`graph_plan_status`), and never report the old run. 4. **`graph_report`** (`graph`, `node_id`, `run_id`, `status`, `report`, optional `pr`, `error`, `outputs`) — `status` is `succeeded`, `failed` or `cancelled`. The response carries `verification` (when judgment is enabled on the tenant), an `outputs` block when the node declares outputs, a `checks` block when its criteria hold `[check]` rows, and `frontier_after`, the nodes runnable after your outcome. They are not claimed; `graph_next_work` claims one. - `node_status: verifying` — the run succeeded under a gate policy and its verdict is pending (a re-judge, or a person's approval). Its dependents are held until the verdict. It is not yours to poll and not yours to retry: move on with `graph_next_work`. 5. **`graph_remember`** (`graph`, `node_id`, `body`, optional `source`) — write a fact about a node at any time: a decision, a measurement, something the next run must know. The server dedupes and can flag a contradiction with an earlier observation; read `judgment` in the response. `graph_plan_status` is the whole plan on one card: counts, frontier, running, verifying, waiting (with `not_before`), failed, unverified runs. Call it when asked how the plan stands, not every turn. ## Following the route When the card's `route` names a profile with a `local_type`, that is the subagent type the router picked for this node. Start a subagent of that type (the Agent tool's `subagent_type`) and hand it the card's `upstream` bundle (the lines to read first), the node's acceptance and the graph and node ids. Because the worker rule holds, the node is claimed by whoever does the work: if you are the worker, you are that subagent's type already, or you run it in your own session; a coordinator never claims on the subagent's behalf. A route without a `local_type` (an `agent` or `a2a` profile) is not something to start locally: work the node yourself and say so. The route is advice, not an order. If you depart from it (a different subagent type, no subagent, or a different approach than the profile implies), write the **override reason** into the report: one line naming the route, what you did instead, and why. Departing silently is the failure; departing with a reason is fine, and the reason is what lets the router learn. ## Evidence: the verdict is what the evidence shows The judge never sees the repository. It asks whether the **evidence** shows each criterion met. Your report is the claim, and a claim is not evidence: a well-written report with nothing behind it scores like no report at all. - **Evidence is captured, not written.** The plugin records every `Bash` command you run, with its exit code and output, and every `Edit`/`Write`, with its path. It attaches the record to `graph_report` for you and replaces anything you put in `evidence`. Reads, greps and globs are not captured. Do not hand-write `evidence`, and do not paste output into the report as proof. - **Prove each line with a command that runs in your own session.** If you claim a test passes, run it. If you claim a file changed, show `git diff --stat`. If you claim a PR is merged, run `gh pr view --json state,mergedAt`. The exit code is part of the record. A failing command is evidence too, and is kept first. - **Long output goes to enforcer-files, whole.** Each output is clipped to 4000 characters, keeping both ends. If `files_base_url` is set in `.enforcer/graph.json` (the older Claude project config is still read), the attach hook uploads any longer output to the user's enforcer-files and the evidence item carries its `file` id; you do nothing. Without that setting, upload the log yourself through Bash, so that the upload is captured too (this is the `files` skill run as a command): `node "${CLAUDE_PLUGIN_ROOT}/bin/files.mjs" upload --dir graph-evidence`. Then cite the file id it prints in the report line the log supports. - **Prose only supports.** The report says which evidence answers which criterion; it cannot stand in for the evidence. At most 20 items are attached: failures first, then the most recent. Run the decisive checks last. - **Cost is stamped for you.** The same hook writes `data.usage` on the run — model, tokens and tool calls since your claim, from your own transcript — so a plan can be costed with `graph_query`. Do not pass it yourself. Output tokens there are an estimate; input and cache counts are exact. - **Keep tool output small; it is re-read every turn.** Each result stays in your context for the rest of the run, so a 7k-token log at turn 10 of 60 is paid for 50 times. Measured: workers spent 96–98% of their input re-reading old tool output. Send long commands to a file and read the end and the failures: `bash scripts/verify.sh > /tmp/v.log 2>&1; echo EXIT=$?; tail -40 /tmp/v.log`, `go test ./... 2>&1 | grep -E '^(FAIL|ok|panic)|--- FAIL' | head -50`. Read a file's range (`sed -n`, `grep -n`, or Read with offset), never `cat` a whole file. Batch independent commands in one call. Do not re-run a gate on a tree nothing changed. Write the report as a numbered list in the card's `criteria` order. Mark each line **MET** or **NOT MET** and name the command or file id that shows it. Report `succeeded` only when every line is met. Otherwise report `failed`, with `error` saying what the next attempt needs. "Not met, because X" scores honestly. ### Worked example ``` # graph_next_work -> node fix-lease-race, run r-81 # criteria: 1. lease race test passes under -race 2. PR merged to main $ go test -race -count=20 ./internal/runs/ # captured: exit 0, output $ gh pr create --fill # captured: the PR URL $ gh pr checks 212 --watch && gh pr merge 212 --squash $ gh pr view 212 --json state,mergedAt # captured: MERGED graph_report status=succeeded pr=https://github.com/o/r/pull/212 report= 1. Race test — MET. `go test -race -count=20 ./internal/runs/` exit 0, "ok example/internal/runs". 2. PR merged — MET. `gh pr view 212`: state MERGED, mergedAt 2026-09-30T02:11Z. ``` Compare a coordinator writing "The subagent reports the race is fixed and the tests pass." That is prose, and it is judged `unsupported_by_evidence` however true it is. ## When the work departs from the plan Acceptance lines go stale: a file moves, an approach fails, an upstream decision changes. Never do something else quietly and report it as met. 1. `graph_remember` on the node: what the line says, what you did instead, and why. 2. In the report, mark that line **NOT MET — STALE: **. Report `failed` unless every line that still applies is met and a person has said the stale one may go. Do not edit the node's acceptance to fit your work; that is the plan author's call. ## When a node fails: remember, remediate, triage A failure means the context was not good enough to pass. The loop fixes the context, or adds the missing work to the plan, rather than stopping. 1. **Remember, then report.** Before any `failed` report, `graph_remember` one structured observation: `Tried: … Blocked by: … Next attempt needs: …`, with `data: {tried, blocked_by, needs: {kind: file|fact|prerequisite|human_decision, what}}`. The next claim card's `observations` carry it. 2. **Remediation launch.** graph-dispatch launches a failed node again (not a DENIED one) by resuming the failed attempt's session, with the previous `error`, the `last_rejection`, and "fix the cause, not the symptom" in the prompt. 3. **Triage.** When the second attempt fails too, the task is not launched a third time. The dispatcher creates a triage node (`data.triage_of`) and runs a deep-tier worker on it. That worker reads the node, its runs, observations and verdicts, then reports its OWN node (never the failed one) with exactly one `data.triage`: - `revise`: the node can't pass as written. The dispatcher patches its description, acceptance or brief with `data.revised_by: triage` and the reason. - `prerequisite`: work is missing. The dispatcher creates the node(s) and a `requires` edge from the failed node to each. - `gate`: the blocker is a human decision or a permission. The dispatcher creates `gate-` naming the decision, adds the `requires` edge that holds the node behind it, and makes no other change. The dispatcher does the writes because a headless worker can't: every hosted write except the graph work loop asks a person. 4. **Bounded.** Each node gets at most 2 task launches and 1 triage per dispatcher session (`--max-attempts`, `--no-triage`). A denial from the environment (a push, a PR, a tool) takes the salvage path instead, never triage. ## Done means merged When the plan has merge nodes (nodes whose acceptance says a PR is merged), an open PR is not done. On a merge node, `succeeded` needs `gh pr view` showing `state` MERGED in your evidence. A PR that is open, failing CI or waiting for review is NOT MET: report `failed` with the PR URL in `pr`. On the node before a merge node, open the PR, pass `pr` and report. Do not merge it unless your own node says to. Push alone: commit in one command, then run `git push -u origin graph/` as its own command, never chained with `&&` or `;` to a test, add or commit (the headless push rule matches only the whole command), then open the PR. To merge, run the plugin's lander, which needs no model in the loop: `"${CLAUDE_PLUGIN_ROOT:?set by the dispatcher; or run `enforcer land `}/bin/land-pr.sh" `. The dispatcher exports `CLAUDE_PLUGIN_ROOT` (Claude Code itself sets it only for hooks and MCP, not your Bash tool); `enforcer land ` is the harness-independent form of the same command. It queues `gh pr merge --squash --auto`, keeps the branch current, and exits 0 merged (printing the evidence below), 2 CI failed, 3 conflicts, 4 timed out. When the base branch requires branches to be up to date, GitHub merges only once the PR is current and green; if it falls behind, `gh pr update-branch ` (or rebase, re-verify, push) and it stays queued. Do not poll `gh pr view` in a loop or re-run the full verify on a branch nothing changed. The merge evidence is the end state, not the steps: `gh pr view --json state,mergeCommit` showing MERGED, and `git merge-base --is-ancestor origin/` exiting 0. ## Coordinating subagents If you fan the plan out, you coordinate and the workers claim. Fan-out is your whole job. 1. Read the frontier with `graph_plan_status`. Spawn at most one worker per frontier node, each in its own worktree. Do not claim anything yourself. 2. Give each worker the graph id, the node key and id you expect it to get, its acceptance lines, and this instruction: *"Load the `enforcer:graph` skill. Claim with `graph_next_work` passing `node: `, work the node you are handed, heartbeat through long steps, and `graph_report` it yourself before you return. Tell me the node key, run id, status and verdict."* **Hand a worker a specific node with `graph_next_work`'s `node` parameter (MCP 0.9.5 or newer), not by claiming by id.** Parallel workers each asking for "the top of the frontier" race and may swap nodes; `node` removes the race and keeps the claim a `graph_next_work` claim the hooks follow. On an older server `node` is unknown: fall back to the guidance above (a plain `graph_next_work`; if it hands back a different node, work that one and say so). Never claim by id through `enforcer_api_write` as a workaround. Also give the worker the card's `route` and `upstream` if you have them; the worker's own claim returns them anyway. 3. When workers return, read `graph_plan_status` again and fan out the new frontier. Never call `graph_report` for a worker, never re-report or re-judge its node, and never heartbeat its run. If a worker's verdict is rejected, a new attempt by a worker fixes it. A better summary from you does not. 4. **Contested resources are not parallel work.** Two frontier nodes with no edge between them can still collide on one shared thing: a migration number, a reservation file, a lockfile, one generated file. The graph only serializes what an edge says. Before fanning out, assign each worker its value (claim the migration number in the repo's reservation file yourself, in its own PR, and put it in the worker's prompt), or add a `requires` edge so the second waits for the first. Never let two workers each discover "the next free one". 5. **Plans written for the `graph-worker` agent have no merge nodes** (see "Writing a plan"): the worker lands its own PR. For an older plan that still has merge nodes, fan out the build work in parallel; land the merges through one worker per repo, in order, each queued with `--auto`. N parallel merge workers on one repo each rebase onto a base the others keep moving, and every move is a full verify for nothing. 6. **Launch each worker at its node's tier.** The card's `model` (when the server sends one) and the node's `data.tier` say how much model it needs; launch the worker at that model (the Agent tool's `model`). The tier maps: `mechanical` → a script (`land-pr.sh`, the release make targets) or Haiku; `standard` → Sonnet; `deep` → Opus. `tier_source` says who decided — `user` beats `planner` beats `rule` beats `jev`, so a tier the user set through MCP is never second-guessed by you. A node with no tier is `standard`, except merge/release/chore (`mechanical`) and a gate or anything touching authz, tokens or a hot-table migration (`deep`). A node on its second attempt goes up one tier. A node marked `detail: thin` has only a pointer for a description: its tier is a rule's guess, so give its worker the design section to read and say the tier is provisional. ## Writing a plan When you author nodes (import, templates, `addNode`), write them so a `graph-worker` agent (this plugin's `agents/graph-worker.md`) can take one cold: - **Start from a template, and let the graph carry HOW.** Create the graph from a template (`software-delivery` or `accounting-close`) or set `node_defaults` per node type (outputs, checks, tier, profile) and graph skills (filtered by node type). The delivery procedure is a skill, not part of the worker: software-delivery attaches `deliver-via-github-pr` (worktree, push, PR, `land-pr.sh`), accounting-close attaches its own. The worker follows the card's skills in order, graph-level first, and with none it only reports what it produced. So a node says WHAT (the outcome and acceptance), never HOW, and does not repeat the defaults. The dispatcher's salvage of a denied push applies only to nodes whose skills include `deliver-via-github-pr`. - **On a software-delivery graph, a code task's acceptance includes "landed".** The worker opens the PR and lands it with `land-pr.sh` inside its own node, so one of the acceptance lines reads like "PR merged to the base branch (merge commit on origin/main)". Do not add a separate merge node per code task: it costs a claim, a rebase and a full verify for nothing, and the hooks capture the merge evidence on the node that did the work. A merge node is only for something no code task owns (a release, a cross-repo cutover). - **`data.tier` is required** on every node: `mechanical`, `standard` or `deep` (see "Coordinating subagents" for what each launches). Add `data.tier_reason` in one line. - **`data.repo` names the repository** (and the base branch, when it is not the default) the node's code lands in. - **A node's description IS its spec.** A worker must be able to start from the description, the acceptance lines and `data.brief` alone. Write the facts in: the files, the numbers, the decisions. Never "see plan section X", "per the plan" or "as in the contract": on one 80-worker graph a 153 KB plan document was re-read by 16 workers and a 189 KB contract by 15, and every later turn paid for them again (orientation was 27% of all tokens; the median first edit came at turn 39.5). Copy the paragraph the node needs into the node. - **`data.brief` limits: at most 5 files and 30 KB in total.** It names the files the worker reads before editing, nothing else. A pointer to a document over 30 KB is banned, even as one of the five: extract the part the node needs into the description, or point at a per-repo `docs/WORKER_BRIEF.md` instead. The plugin's `test/check_skill.py plan` rejects a sample plan that breaks these. - **`scout` nodes write the repo brief.** For a repo that has no `docs/WORKER_BRIEF.md`, add one node of `type: scout` (with `data.repo`) before the repo's code nodes, and make them depend on it. Its job is to write and land `docs/WORKER_BRIEF.md` for that repo, at most 8 KB, with these sections: layout (where things live), verify loop (the exact commands), traps (each with its real error message), where to look for X, and docs never to read whole (name them and their size). Its acceptance: the file exists, is <= 8 KB, has those five sections, and is merged. Workers read `CLAUDE.md` and `docs/WORKER_BRIEF.md` first, so the brief replaces their exploration. - **`data.max_turns` is optional.** Workers have an 80-turn cap; a deep node may raise it, and the dispatcher passes it on. - Contested resources (a migration number, a reservation file) go in the node description, assigned by you, never discovered by the worker. ## Inputs and outputs - **Read inputs from the card; never re-derive them.** `inputs` says where each value came from (`image (docker_image) = ghcr.io/… — from ship-fix attempt 1, extracted 0.97`); `input_values` holds it exactly. Do not re-run `git rev-parse`, re-read a log or re-compute a number to get a value the card already handed you: the point is that the value is the one the upstream run was judged on. An input marked `ABSENT` is not an error — proceed without it, or report why you cannot. **Inputs never cross epochs:** after a reset they resolve from the current epoch's runs only, unless the plan writes the input as `.@previous`. Don't go looking for an earlier epoch's value. - **Declare `outputs` only for computed values.** Anything verbatim in your evidence — a PR URL, a commit sha, a pushed image — is extracted from the evidence without you. Pass `outputs: {name: value}` on `graph_report` only for what you computed or chose (a count, a variance, an id you picked). Each declared value is judged against the evidence and dropped when the evidence does not show it, so the command that produced it must have run. - **A verifying node is not yours to poll.** Its verdict arrives on the run (a worker re-judge, or a person approving it in `graph_review`). Do not heartbeat it, re-report it or call `graph_plan_status` in a loop waiting for it; take other work, or stop. - **A reset is a human's call.** `graph_reset` starts a new epoch and can cancel running work. Never reset a graph to get unstuck, to retry, or because a loop looks due (the server's loop worker does that). Only reset when a person asks, through the tool's own confirmation. ## Two conventions that silently invert the plan if reversed - **A dependency edge points FROM the dependent TO its prerequisite.** `build --requires--> setup` means "build requires setup". The frontier is the set of unfinished nodes whose every outgoing `requires` edge points at a done node. Reverse an edge and the plan runs backwards without an error. - **A repeated step is a node with several runs, never a loop edge.** A retry is attempt N+1 on the same node. On a dag-mode graph an edge that would close a cycle is refused with `409 edge_would_create_cycle`. ## Before stopping Report the run, or say explicitly that you are leaving it open. In a top-level session the plugin's Stop hook sends you back once if a run is open and unreported. A subagent gets no such reminder.