{ "schema": "chi-bench/submission/v1", "submission": { "id": "erius-opus5", "team": "Michael Johnson (MJ)", "contact": "johnson.michael@gmail.com", "agent": "erius", "model": "anthropic/claude-opus-5", "notes": "Erius on Opus 5: one GAMPO-governed agent (gampo-advisory), one model (claude-opus-5),\nstrict single-attempt pass@1 (n_attempts = 1), run via the Claude Max subscription (CLI\nauth). Judge pinned to claude-opus-4-7 per leaderboard policy; agent and judge run on Max,\nand the ONLY component that bills the funded API key is the care-management patient\nsimulator (sonnet-4-5), which the benchmark structurally requires.\n\nGovernance is an ANSWER-BLIND, per-task definition-of-done: a specification keyed only to\neach case's own visible policy and to published United States standards (CMS-0057-F / Da\nVinci PAS, NCD/LCD and NASS/InterQual criteria, CMS/AMA coding read against the chart, and\nthe CCM/APCM/GUIDE care-management programs), never the hidden answer key, rubric, or\nsolution.\n\nPer-domain configuration (all single-attempt, same agent and model):\n - Prior-authorization (pa_provider): answer-blind advisory with chart-documentation-\n fidelity facts (place-of-service and diagnosis-code corrections traced to the case's\n own chart, never the gold), default reasoning effort. 18/25 = 72%.\n - Utilization-management (pa_um): answer-blind advisory, xhigh reasoning effort. A\n terminal sign-off commit lever was tested and REJECTED under a zero-regression rule\n (it recovered one sign-off task but regressed a passing one, net 9/25), so the filed\n UM slice is the plain advisory board. 9/25 = 36%.\n - Care-management (cm): the GAMPO procedure only, under the same gampo-advisory agent\n with no content-advisory file injected (the procedure beats the CMS content advisory\n on this domain), max reasoning effort, sonnet-4-5 patient simulator. 14/25 = 56%.\n\nOverall: 41/75 = 54.7% strict single-attempt pass@1. Exploratory: single trial per task,\n25 tasks per domain. This is a cross-generation replication of the Opus 4.8 erius result\n(filed separately at 37.3% under the generic procedure); it is not a resubmission of that\nentry.\n", "submitted_at": "2026-07-27T03:15:17Z" }, "dataset": { "version": "chi-bench-v1.0.0", "domains": [ "pa_provider", "pa_um", "cm" ], "name": "chi-bench" }, "results": { "overall": { "n_trials": 75, "n_tasks": 75, "pass_at_1": 0.5466666666666666, "mean_cost_usd": 0.0, "mean_walltime_s": 0.0 }, "per_domain": { "pa_provider": { "n_trials": 25, "n_tasks": 25, "pass_at_1": 0.72, "mean_cost_usd": 0.0, "mean_walltime_s": 0.0 }, "pa_um": { "n_trials": 25, "n_tasks": 25, "pass_at_1": 0.36, "mean_cost_usd": 0.0, "mean_walltime_s": 0.0 }, "cm": { "n_trials": 25, "n_tasks": 25, "pass_at_1": 0.56, "mean_cost_usd": 0.0, "mean_walltime_s": 0.0 } }, "mean_cost_usd": 0.0, "mean_walltime_s": 0.0 }, "provenance": { "chi_bench_git_sha": null, "image_digest": null, "judge_model": "claude-opus-4-7", "harness_version": "0.1.0" } }