--- name: arena description: "Spawn N parallel candidates at the same task, pick a base, and graft the strongest parts of the others into it. Use for /arena, 'arena this', 'throw it in the arena', competing approaches, or when one attempt at a non-trivial artifact would lock in the wrong shape." license: MIT metadata: adapted-from: "cursor/plugins/pstack/skills/arena" version: "1.0.0" --- # Arena Fan out N parallel attempts at the same task. Read every candidate end to end. Pick the strongest as the base. Graft the best ideas from the others into it. Verify the synthesized result. Open a todo list with one entry per phase: Frame, Fan out, Cross-judge, Pick, Graft, Verify. ## Phase A: Frame The candidates all receive the same prompt, so the prompt is the contract. 1. State the artifact each candidate produces. 2. Derive a rubric: what success looks like for this task, turned into 3-6 concrete gradeable criteria. Candidates see the task; the picker uses the rubric. 3. Choose the runners. Prefer diverse reasoning where judgment matters; use identical runners when the work is generation-bound. Route by role; Kiro selects the model. 4. Give each candidate its own writable output location (a git worktree where possible, otherwise a separate scratch dir) so they never write the same path (Separate Before Serializing Shared State). ## Phase B: Fan out Spawn all N candidates in parallel via `invoke_sub_agent`, each with the task, the shared grounding, its own output path, and instructions to produce both the artifact and a short rationale naming the alternatives it considered and rejected. If a candidate fails, proceed with N-1 and note the dropout. ## Phase C: Cross-judge Spawn one read-only judge (prefer a different reasoning profile from the parent) that sees the rubric and the candidates, scores each criterion, and recommends a base with rationale. ## Phase D: Pick a base Read every candidate end to end before picking. Score against the rubric criterion by criterion, not on holistic feel. Compare with the cross-judge; agreement confirms, disagreement means bias or an ambiguous rubric. Pick the base a future maintainer can extend most easily without breaking invariants; prefer the cleaner boundary or smaller API on a tie. Record the pick and reason. ## Phase E: Graft Walk each losing candidate once more and identify what is worth porting (usually one or two things each). Fold each graft in by hand so the result stays coherent under one mental model. Record what was grafted, from where, and what was rejected and why. If candidates converged on the same shape, ship the consensus and skip grafting. If they wildly diverged, Phase A was under-specified: reframe and re-run rather than averaging. ## Phase F: Verify The synthesized artifact holds up under the same scrutiny as any other output (Prove It Works). If verification surfaces a problem the arena missed, either reframe (Phase A was wrong) or redo the graft. ## Outputs One synthesized artifact plus a short synthesis note naming the base, the grafts and their sources, the rejections, any dropouts, and the verification result.