--- name: harness-audit description: >- Audit an agent's harness against the four rings: containment, guides, sensors, permissions. Use when the user asks whether their agent setup is safe, wants a review of sandboxing / permissions / hooks / AGENTS.md guardrails, or asks what a bad session could break. Read-only: reports the blast radius and a hardening list, changes nothing. Do NOT use for context layout (context-audit) or cost routing (gate-check). --- # Audit the four rings Theory: [The four rings](https://secondbrainos.dev/#/read/course-4-harness/the-four-rings) and [Harness practice](https://secondbrainos.dev/#/read/course-4-harness/harness-practice). Build from the outside in: containment, guides, sensors, permissions. The order matters because each ring must hold when every ring inside it fails — a guide can be ignored, a sensor can miss, an approval can be misclicked; a wall does not care. ## Core rule Report and rank; never harden anything yourself. The audit's one organising question is the blast radius: describe the worst session possible under the current setup and what it costs. If the user asks you to apply fixes afterwards, that is a normal edit session, not this skill. ## Workflow 1. **Ring one — containment: what can it physically reach?** Where does the agent run: the user's checkout or a worktree/container? Which credentials are in its environment — a read-only database user or the application's? Open internet or an allowlist? Real API keys or test keys? The finding is one sentence: "the worst session deletes X and spends Y." 2. **Ring two — guides: what does it read before acting?** Find CLAUDE.md / AGENTS.md and tool descriptions. Every line should be a rule the agent can act on — flag philosophy, history, and anything over roughly a page. Flag missing rules for the repo's real landmines (migrations, generated files, deploy scripts). 3. **Ring three — sensors: what fires when it errs?** Is there one command (`make check` or equivalent) that runs tests, linter and types? Does the harness run it automatically after edits — a hook — or does correctness depend on the model remembering? A sensor the model must invoke voluntarily is a guide, not a sensor. 4. **Ring four — permissions: what still needs a human?** List what is auto-approved. The irreversible set — push to main, deploy, spend, delete data, send messages — should need a person; everything reversible inside the walls should not. Flag both failure modes: irreversible actions auto-approved, and reversible ones prompting so often the user rubber-stamps (approval fatigue is a hole in this ring, not a virtue). 5. **Check the order.** The classic smell is an inner ring doing an outer ring's job: a prompt line saying "never touch prod" where a credential should have been removed. Rules in guides that could be walls belong in ring one. ## Output format ``` Harness audit — blast radius: ring 1 containment ring 2 guides ring 3 sensors ring 4 permissions Hardening list, outermost first: 1. ... ``` Rank fixes outside-in — a wall before a better prompt, a sensor before a politer rule. Close by restating the blast radius as it would be after fix 1 alone.