--- name: troubleshoot description: | Diagnose and plan fixes for errors/bugs with Codex-first multi-agent collaboration (Codex + Opus 4.6 + Agent Teams). Codex CLI is consulted in EVERY phase for deep code reasoning, hypothesis evaluation, and fix validation. Phase 1: Error reproduction & context gathering (Opus subagent 1M context + Codex initial analysis + Claude user interaction). Phase 2: Parallel diagnosis (Agent Teams: Root Cause Analyst [Codex-driven] + Impact Investigator [Opus + Codex risk analysis]). Phase 3: Fix plan synthesis, Codex validation & user approval. Fix implementation is handled separately by /team-execute. metadata: short-description: Codex-first error/bug diagnosis with Agent Teams (Diagnosis phase) --- # Troubleshoot **Codex-first error/bug diagnosis skill leveraging Codex deep reasoning, Opus 1M context, and Agent Teams.** > Preflight: ensure codex CLI is current (see codex-system skill). ## Overview This skill handles the diagnosis phases (Phase 1-3) with a **Codex-first approach**: Codex CLI is consulted proactively in every phase for pattern recognition, hypothesis evaluation, root cause reasoning, and fix validation. Fix implementation and review are done via `/team-execute`. ``` /troubleshoot <- This skill (diagnosis & fix planning) | After approval /team-execute <- Parallel fix implementation (Phase 1) | After completion Phase 2 REVIEW <- Parallel review (regression check) ``` ## Workflow ``` Phase 1: REPRODUCE & UNDERSTAND (Opus 1M context + Codex Initial Analysis + Claude Lead) Opus subagent analyzes the error context, Codex generates initial hypotheses, Claude gathers details from the user | Phase 2: DIAGNOSE (Agent Teams -- Parallel, Codex-driven) Root Cause Analyst (Codex mandatory) <-> Impact Investigator (Opus + Codex) communicate bidirectionally Both teammates consult Codex for deep reasoning throughout analysis | Phase 3: FIX PLAN & APPROVE (Codex Validation + Claude Lead + User) Integrate diagnosis results, validate fix plan with Codex, get user approval ``` --- ## Phase 1: REPRODUCE & UNDERSTAND (Opus Subagent + Codex + Claude Lead) **Reproduce the error and gather full context with Opus subagent's 1M context, then consult Codex for initial hypothesis generation, while Claude interacts with the user.** > Main orchestrator context is precious. Large-scale error context analysis is delegated to Opus subagent (1M context). > Codex is consulted early for pattern recognition and hypothesis generation. ### Step 0: Resolve Workspace Resolve this bug's deterministic workspace once. The title becomes file and directory names, so give it a short English descriptor of the bug -- not the user's raw wording, which the Language Protocol keeps out of paths: ```bash python3 .claude/skills/_shared/workspace.py --skill troubleshoot --title "{short English title}" --create ``` This prints one JSON object: `slug`, `team_name`, and `paths` (`bug_report`, `context`, `root_cause`, `impact`, `diagnosis`, `state_input`, `team_dir`). Exit 0 resolved/created; 1 bad args; 2 applies only to `--verify` (used later in Phase 3); 3 the workspace directories could not be created. Use `{slug}`, `{team_name}`, and every `paths.*` value from this JSON verbatim for the rest of this skill -- do not re-derive them by hand in a later phase. ### Step 1: Gather Error Details from User Ask the user to provide: 1. **Error message / stack trace**: Full error output 2. **Reproduction steps**: How to trigger the error 3. **Expected vs actual behavior**: What should happen vs what happens 4. **Environment**: OS, Python version, dependency versions 5. **Recent changes**: What changed before the error appeared (if known) ### Step 2: Reproduce & Capture Context (repro.py) First run the bundled script for the **mechanical** capture — it runs the failing command under a deadline, records stdout/stderr/exit code + extracted traceback to a log file keyed by `--label`, and gathers recent git history (plus optional last-commit context for a stack-trace file): ```bash python3 .claude/skills/troubleshoot/repro.py "" \ --label {slug}-initial [--file ] [--timeout 120] ``` Always pass `--label {slug}-initial` here: the log path is `.claude/logs/troubleshoot-repro-{label}.log` and an unlabelled run reuses one shared file, so the Phase 3 fix-verification run (Step 2 task 3) would otherwise overwrite the original failure evidence this whole diagnosis rests on. Exit codes: `0` capture completed; `1` bad arguments (including an unusable `--label` or `--bisect-good` ref, checked *before* the command runs); `2` the observed exit code differs from `--expect-exit` (not used in Phase 1); `3` the repro command timed out or the log could not be written. A failing repro command is the expected case and is still exit `0` — its result is the JSON `exit_code`. Read the JSON fields: `exit_code`, `timed_out`, `stdout_tail`, `stderr_tail`, `traceback`, `traceback_format`, `git_available`, `git_error`, `recent_commits`, `blame`, `blame_error`, `bisect`, `log_file`, `artifacts`. Two fields exist to stop a null being over-read: `traceback` is only extracted for CPython tracebacks (`traceback_format: "python"`), so `null` there means "no *Python* traceback" — a Node/Go/pytest-assertion stack is in `stderr_tail`. And `git_available: false` with a `git_error` means history could not be read at all; that is not the same as "no relevant recent history". On `timed_out: true` (exit 3) the command has no usable result: raise the `--timeout`, narrow the repro command, or treat the hang itself as the bug — do not proceed as if the capture succeeded. To scope a regression, add `--bisect-good `. It reports the `bisect` object (`candidate_commits`, `candidate_count`, `path_filter`, `bisect_command`) — the commits an actual `git bisect` would search, plus the command to start it. The script never checks out a commit itself, so driving the bisect stays the Impact Investigator's call in Phase 2. Then hand that captured context to `general-purpose-opus` for the **judgment** part — do NOT re-run the command or re-fetch git history: ``` Task tool: subagent_type: "general-purpose-opus" prompt: | Analyze this reproduced error (already captured by repro.py): Error: {error message / stack trace} repro.py JSON: {exit_code, traceback, recent_commits, log_file} Tasks: 1. Read all files mentioned in the traceback; trace the execution flow leading to the error and identify the immediate cause (what line fails). 2. Look for related tests and whether they pass/fail. 3. Check if similar patterns exist elsewhere in the codebase. Use Glob, Grep, and Read tools to investigate thoroughly. Save analysis to `{paths.context}` (from Phase 1 Step 0). Return concise summary (5-7 key findings). ``` ### Step 2.5: Codex Initial Error Pattern Analysis Consult Codex for initial hypothesis generation before creating the Bug Report. Write the prompt to a file, then invoke the wrapper: ```text Objective: Analyze this error and generate initial hypotheses for root cause. Context: - Error: {error message / stack trace} - Failing location: {file:line from Opus subagent analysis} - Execution flow: {call chain from Opus subagent analysis} Constraints: - Focus on root cause categories (state mutation, boundary, concurrency, dependency, type/contract) - Rank hypotheses by likelihood - Suggest specific code areas to investigate for each hypothesis Output format: ## Error Pattern Recognition ## Hypotheses (ranked by likelihood) ## Investigation Plan (per hypothesis) ## Known Similar Patterns ``` ```bash python3 .claude/skills/_shared/codex_consult.py --prompt-file .claude/logs/codex/prompt-troubleshoot-initial.md --label troubleshoot-initial ``` `.claude/skills/_shared/codex_consult.py` exits 0 when Codex answered normally, 2 if the Codex CLI is not installed, 3 if Codex failed or timed out -- check the JSON `ok` field and read `response_file` for the answer (`error`/`stderr_file` explain a failure). Every later Codex consultation in this skill follows this same write-prompt-then-invoke pattern without repeating these exit codes. Use Codex's analysis to strengthen the Initial Hypotheses section of the Bug Report. ### Step 3: Create Bug Report Combine error details + codebase analysis + Codex initial hypotheses into a Bug Report following the template contract in `references/bug-report-template.md`. Save it to `{paths.bug_report}` (from Step 0), then validate it: ```bash python3 .claude/skills/_shared/validate_doc.py --contract bug-report --file {paths.bug_report} ``` `references/bug-report-template.md` is the single source of truth for the required sections; the `bug-report` contract is pinned to that template by `tests/test_validate_doc.py`. Do not work from a section list retyped here -- that drift is exactly what this fix removed. Exit 0 means every required section is present; exit 2 means one is missing, and the JSON `sections_missing` names it. Fill the gap before proceeding; exit 1 means the file does not exist. Both Phase 2 teammates read this file, and Phase 3's `--verify` gate requires it, so it must exist on disk -- not only in this conversation. --- ## Phase 2: DIAGNOSE (Agent Teams — Parallel) **Launch Root Cause Analyst and Impact Investigator in parallel via Agent Teams with bidirectional communication. Both teammates MUST consult Codex for deep reasoning tasks.** > Key difference from subagents: Teammates can communicate with each other. > Root Cause Analyst's findings change Impact Investigator's scope, and Impact Investigator's context informs root cause analysis. ### Team Setup ``` Create an agent team named `{team_name}` for troubleshooting: {slug} Spawn two teammates: 1. **Root Cause Analyst** — Uses Codex CLI as PRIMARY analysis engine for deep code reasoning Prompt: "You are the Root Cause Analyst for bug: {slug}. Your job: Identify the definitive root cause of this error through deep code analysis. Codex CLI is your PRIMARY tool for reasoning about code behavior. Bug Report: read `{paths.bug_report}` (written and validated in Phase 1 Step 3). Tasks: 1. Trace the execution flow step by step from entry point to error 2. Evaluate each hypothesis from the Bug Report: - Gather evidence FOR and AGAINST each hypothesis - Eliminate hypotheses that contradict the evidence 3. Identify the root cause (not just the symptom): - What is the underlying defect? - Why does it manifest as this specific error? - Under what conditions does it trigger? 4. Propose fix approaches (at least 2 alternatives): - Approach A: {description, pros, cons} - Approach B: {description, pros, cons} - Recommended approach with rationale ## Codex Analysis Protocol (MANDATORY) You MUST consult Codex for EACH of the following analysis tasks. Do NOT skip Codex consultation — it is the primary reasoning engine for this role. Each consultation below follows the same shape: write the prompt to a file, then run `python3 .claude/skills/_shared/codex_consult.py --prompt-file --label