--- name: interrogate description: "Multiple LLM reviewers challenge changes from independent angles. Use before any PR is opened or integrated (run by the change's owner, never by a worker on its own unit), on any contested design decision, or for \"interrogate\", \"adversarial review\", \"multi-model review\", \"challenge this\", \"stress test this code\", \"find blind spots\", or \"tear this apart\"." --- # Interrogate On Codex, spawning a reviewer is `spawn_agent`; see `../poteto-mode/references/codex-tools.md`. Spawn one reviewer per configured model to adversarially review code changes. Each model gets the same prompt and rubric. The adversarial signal comes from model diversity, not assigned personas. Models differ in blind spots, priors, and reasoning patterns. Agreement across models is high-confidence signal; lone-model findings are worth reading but lower confidence. The deliverable is a synthesized verdict. Do NOT auto-apply changes. Owner only. The main thread or the `orchestrator` that owns the change runs this. A worker (`developer`, `developer-codex`, `tester`, and the rest) never runs it on its own unit; it reports, and its owner runs it. Block only on material risk or missing evidence, never on preference. Check scope alongside correctness: scope creep, missing acceptance criteria, or architecture drift are findings in their own right, not just bugs. ## Step 1, Determine Scope Identify what to review from context: - If the user points at specific files or a diff, use that - If on a feature branch, run `git diff main...HEAD` (or the appropriate base branch) for the full changeset - If the user's message references recent work, gather the relevant files Package the diff (or file contents) plus any surrounding context files the reviewers need to understand the code. ## Step 2, State the Intent Before spawning reviewers, state the intent explicitly. What is this code trying to accomplish? Derive this from: - The user's message - Commit messages - PR description if one exists - The code itself Write one clear paragraph. Reviewers challenge whether the work achieves the intent well, not whether the intent itself is correct. If you're unsure about the intent, ask the user before proceeding. ## Step 3, Spawn Reviewers Spawn every reviewer at once, one `reviewer` per entry in the `interrogate reviewers` row of `plugins/pstack-nikki/models.json`. Label them Reviewer A, B, C... in row order. Row absent -> two reviewers, `opus` and `sonnet`. Each brief: - pins `model` to its entry (entry `inherit` -> omit `model`) - FORBIDDEN: no writes, no commits, inspection commands only - carries the same filled template, so every model applies the same lens Read `references/reviewer-prompt.md` and fill in the template with: 1. The stated intent 2. The diff or file contents 3. The review rubric from `references/rubric.md` 4. The code-quality lens from `references/code-quality-review.md` The same filled template goes to all reviewers, so every model applies the code-quality lens. Each reviewer produces structured findings as described in the prompt template. ## Step 4, Synthesize As results come back, build a unified picture: 1. **Parse all findings** from the reviewers 2. **Identify consensus**. Findings raised by 2+ models independently are highest signal. 3. **Identify lone-model findings**. Still worth reading, but weight accordingly. 4. **Deduplicate**. Different models may describe the same issue differently. Merge these and note which models raised it. 5. **Note disagreements**. If one model flags something and another explicitly says the opposite, that's useful context for the verdict. ## Step 5, Lead Judgment You are the lead reviewer, a pragmatic senior engineer, not a neutral aggregator. Read `references/lead-judgment.md` for the full framework. Reviewers only see a slice of the codebase. You have the full context (the goal, the constraints, the timeline, which tradeoffs were already considered). Use that context aggressively. Sort every finding into one of the four buckets defined in Output Format below. For each, record which model(s) raised it and a one-line rationale for the categorization. ## Output Format Present the verdict in this structure: ### Intent > [The stated intent paragraph from Step 2] ### Reviewers - Reviewer [label]: [model name], [N findings] (one bullet per reviewer) ### Act On [Real issues in correctness, security, or maintainability given the actual goals: these would block a real PR. Each: description, which models raised it, why it matters.] ### Consider [Legitimate, but you are unsure the fix outweighs its cost right now. Each: description, which models raised it, the tradeoff.] ### Noted [Technically valid, not actionable: context-dependent, premature optimization, or low-impact at this stage. Brief list.] ### Dismissed [Wrong, nitpicky, or missing context. Brief rationale each, so the user can override your judgment if they disagree.] ### Agreement Map [Where did models agree, where did they diverge, and what does the pattern of agreement/disagreement tell us?] ## Models Stamped from `plugins/pstack-nikki/models.json` (edit there, rerun `generate-models.py`). Row absent -> omit `model`, child inherits. A spawner reads the entry for its own harness. - `interrogate reviewers`: On Claude Code: `opus`, `sonnet`. On Codex: `gpt-5.6-sol`, `gpt-5.6-terra`. On Copilot CLI: `claude-opus-5`, `gpt-5.5`, `claude-sonnet-5`.