# TRACE Verifier Stress Tests Full code and artifacts for stress-testing automated verifiers in long-horizon tool-using agent workflows. TRACE asks a narrow question: - When an agent is optimized against an automated verifier, is the verifier - measuring grounded task completion, or is it exposing a shortcut the agent can learn to exploit? This repository is actively being developed, and currently contains two pieces of the work: 1. A synthetic ResearchOps environment for controlled verifier stress tests. 2. tau2-bench analysis utilities and saved public benchmark diagnostics used for the Agentic AI Summit poster. ## Install ```bash python -m venv .venv source .venv/bin/activate pip install -e ".[dev]" ``` ## Synthetic TRACE Study Run a small smoke test: ```bash pytest -q python scripts/run_toy_rollouts.py --n-tasks 3 --n-rollouts 2 python scripts/analyze_rollouts.py --input outputs/toy_rollouts.jsonl ``` Run the deterministic stress-test grid: ```bash python scripts/run_stress_tests.py --n-tasks 25 --n-rollouts 3 --no-ray python scripts/analyze_stress_tests.py --input outputs/stress_tests.jsonl ``` The saved poster snapshot is in `outputs/poster/`. ## tau2-bench Diagnostic The saved tau2 run analyzes 164 native tau2 LLM trajectories: - retail base split: 114 tasks - airline base split: 50 tasks - model used for the saved run: `gpt-4.1-mini` Example analysis: > All 164 runs produced final assistant text, but only 79/164 passed the local > official reward threshold. The remaining 85/164 runs show why final-answer > checks can overestimate grounded workflow success. Rebuild the audit tables from the saved analysis artifacts: ```bash python scripts/audit_taubench_llm_results.py \ --analysis-dir outputs/taubench_llm_analysis/20260728_tau2_fullbase_gpt41mini_164_analysis \ --manifest outputs/taubench_llm_runs/20260728_tau2_fullbase_gpt41mini_164/manifest.json \ --run-prefix 20260728_tau2_fullbase_gpt41mini_164 ``` To rerun tau2 trajectories, install tau2-bench separately and set: ```bash export TAU2_BENCH_DIR=/path/to/tau2-bench # Set OPENAI_API_KEY in your shell before launching a paid LLM rerun. python scripts/run_taubench_llm_overnight.py \ --task-source split \ --task-split base \ --n-tasks-per-domain all \ --domains retail airline \ --parallel \ --max-concurrency 2 \ --max-steps 100 \ --timeout 1200 \ --num-trials 1 \ --run-prefix 20260728_tau2_fullbase_gpt41mini_164 ``` This is a public benchmark diagnostic, not a tau-bench leaderboard submission. ## Repository Map ```text src/trace_rl/ Synthetic TRACE environment and reward code environments/trace_researchops/ OpenEnv-style packaged environment variant scripts/ Rollout, analysis, and audit scripts tests/ Unit tests for the synthetic environment outputs/poster/ Saved synthetic TRACE poster artifacts outputs/taubench_llm_analysis/ Saved tau2 analysis and audit tables outputs/taubench_llm_raw/ Saved raw tau2 result JSON files poster/generated/ Poster-ready generated tables ``` ## Scope The artifacts here support reproducible public claims about synthetic verifier stress tests and public tau2 trajectory diagnostics. They do not claim production failure rates, confidential enterprise results, or model leaderboard performance.