VeriHarness

Scaling Agentic Verification for Long-Horizon Tasks

Same model. Better evidence.
Check what agent outputs disagree on, and what they all get wrong. Deliver a better artifact with a record of why.

arXiv Website Dataset License

VeriHarness is a verification harness for long-horizon general agents. Given a task and *N* independent rollouts of the same agent on it, VeriHarness uses **the same model that generated the rollouts** to check them against the task environment and to deliver one artifact that is better than a typical rollout, together with a record of the evidence behind every change. The verifier is not a stronger judge. Its advantage comes from two sources only: the structure of the rollout pool, and evidence it actively acquires from the environment (files, data, recomputable quantities). ## Overview and results
N rollouts go to a disagreement resolver and a consensus challenger, run in parallel with environment evidence; adjudication selects or reconstructs the artifact Bar charts on five benchmarks: single rollout, the best prior LLM-as-a-verifier, and VeriHarness, with VeriHarness highest on each
*Left:* the verifier is the generator's own model inside a harness. A disagreement resolver tests the claims on which the *N* rollouts differ, and a consensus challenger seeks evidence against the claims they share; adjudication combines their findings to select and revise the artifact. *Right:* with Gemini 3.5 Flash as both generator and verifier, VeriHarness against a single rollout and the best prior [LLM-as-a-Verifier](https://arxiv.org/abs/2607.05391) on the five benchmarks. Across the five benchmarks and two frontier models it achieves the highest selection scores among the evaluated baselines, and evidence-backed revision raises the average gain over a single rollout to 6.2 points with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. The full pool of about 26,000 rollouts behind these numbers is released (see [Data](#data)). ## Contents - [Overview and results](#overview-and-results) - [How it works](#how-it-works) - [Repository layout](#repository-layout) - [Setup](#setup) - [Usage](#usage) - [Benchmarks and grading](#benchmarks-and-grading) - [Data](#data) - [Isolation](#isolation) - [Native environments (optional)](#native-environments-optional) - [Operating notes](#operating-notes) - [Skills](#skills) - [Citation](#citation) ## How it works

Worked example: three candidate rollouts report FY2025 revenue; the resolver settles which source applies (draft vs final report), the challenger tests the unit every rollout shared (USD vs EUR); adjudication reviews both records and delivers a corrected report with a verification record

Rollouts of one model either disagree on a claim or agree on it, and the two cases call for different checks: * **Disagreement → resolve.** The alternatives are already on the table. The *resolver* picks the check that best separates them, runs it against the environment, and eliminates the candidates the evidence contradicts. * **Consensus → challenge.** Agreement is not evidence: rollouts share a model and therefore blind spots. The *challenger* looks for ways a shared value, a shared reading, or a shared omission could be wrong, and tests them. Each task runs four model turns in three mutually isolated sessions: ``` materialized task ──► resolver (own session) ──► ledger_elim.json └──► challenger (own session) ──► ledger_fals.json │ (run concurrently; neither sees the other) ▼ adjudication (fresh session: both records + task + rollouts) │ └─► finish.json {base, work[], open[]} ▼ delivery (same session) ──► out/deliverables/ + repair.json ``` Adjudication names a **base** rollout, an evidence-backed **revision plan** (`work`), and the claims the evidence left **unresolved** (`open`). Selection, revision and reconstruction are one knob: an empty plan returns the base unchanged, a non-empty plan revises it, and `base: "none"` rebuilds the artifact from the task inputs. The program only materializes workspaces, routes files between sessions and collects outputs. Which claims to examine, which checks to run and when to stop are the model's decisions. The single programmatic gate is content-blind: a delivered bundle must contain every file of its base under the same name. ## Repository layout ``` harness/ driver.py one task: the four turns above runner.py many tasks: global work pool with per-model and per-benchmark caps; resumable score.py score a run against the archived pool (selection and final scores) config.py locations and model lanes; all host specifics come from the environment views.py plain-text views rendered beside binary artifacts (.cells.tsv, .text.txt) prompts/ CHARTER (system prompt), MISSION, the two investigation playbooks and record formats, ADJUDICATE, REPAIR skills/ the skill library (evidence-*, resolve-*, falsify-*, repair-*) materialize/ adapters that turn a benchmark's archived rollouts into task workspaces grade/ wrappers around each benchmark's own grader, for re-grading revised artifacts env/ optional native execution environments (task images) and the office/docs images scripts/ mount-namespace jail, agent-runtime setup, benchmark setup, local model proxy benchmarks/ our grading configuration and credential templates per benchmark pi-home/ the agent runtime's model registry (models.json) ``` ## Setup Requirements: Linux with unprivileged user namespaces (for the jail), Python 3.10+, Node.js 22.19+ and npm (for the agent runtime), git, curl, `uv` (for the benchmark checkouts), and Docker for the graders and the optional native environments. ```bash pip install -r requirements.txt harness/scripts/setup_pi.sh # installs the pinned pi agent runtime into harness/vendor/ ``` The harness drives the open-source [pi](https://www.npmjs.com/package/@earendil-works/pi-coding-agent) coding-agent runtime and works with any model provider pi supports (Anthropic, OpenAI, Google AI Studio, Vertex AI, OpenAI-compatible endpoints and others); no Google account or service is required. A *lane* is the verifier model; by default a pool is verified by the model that generated it. The two lanes used in the paper are defined in `harness/config.py` (`LANES`); edit them or pass `--provider`/`--model` to the driver to use other models, for example `--provider anthropic --model claude-opus-4-8` with `ANTHROPIC_API_KEY` set. * **Gemini on Vertex AI** (`flash` lane): pi's `google-vertex` provider with application-default credentials (`gcloud auth application-default login`) and `GOOGLE_CLOUD_PROJECT` set. The jail exposes `~/.config/gcloud` read-only for this purpose. * **Claude on Vertex AI** (`opus` lane): served through a local [litellm](https://github.com/BerriAI/litellm) proxy that the runner starts on demand (`harness/scripts/litellm_up.sh`, port 4180; the registry entry is `vertex-litellm` in `harness/pi-home/models.json`): ```bash harness/scripts/setup_litellm.sh export VERTEX_PROJECT= # and optionally VERTEX_LOCATION ``` Host-specific locations are environment variables (see `harness/config.py`): | Variable | Meaning | Default | | --------------------------------------------------------------- | --------------------------------------------------------------------------------- | --------------------------------------------------- | | `VERIHARNESS_DATA` | materialized rollout pools | `./data` | | `VERIHARNESS_RUNS` | run outputs | `./runs` | | `VERIHARNESS_BENCH_ROOT` | upstream benchmark checkouts and archived rollouts (materialize and grade only) | unset | | `VERIHARNESS_TMP` | staging directory for the graders' containers (bind-mounted, so a real directory) | `/var/tmp` | | `VERIHARNESS_WB_INDEX` | index of archived WorkBuddy run directories (materialize only) | `/benchmarks/workbuddy/wb_index.json` | | `VERIHARNESS_IMAGE_`, `VERIHARNESS_IMAGE_SB2_GRADER` | image overrides (see "Native environments" and`harness/grade/sb2.py`) | per bench | `/_worlds/` holds the task environments that accompany the pools: the APEX world archives and the WorkBuddy task repositories with their image markers. The materializers link into it, grading and native mode read it, and the jail exposes it read-only. It is produced together with the pools and is not part of this repository. ## Usage **Task workspace.** Everything the verifier sees is a directory: ``` ///tasks// spec/task.md the task description workspace/ the task's input files rollouts/rNN/deliverables/ what rollout NN delivered rollouts/rNN/trajectory/ its execution trace, if available ///meta.json archived per-rollout scores; never visible to the verifier ``` `python3 -m harness.materialize ` builds these from an archive of rollouts. The adapters encode the archive layout our pools were generated into; to verify your own rollouts, produce the layout above directly or write an adapter that yields `materialize.base.Task` objects. **Run one task.** ```bash python3 -m harness.driver --provider google-vertex --model gemini-3.5-flash --thinking high ``` The driver writes the two investigation records (`elim/`, `fals/`, `ledger_*.json`), `finish.json`, `repair.json`, `driver.log`, the model transcripts under `session/` and the result under `out/deliverables/` into the task directory. `--contract pick-only` stops after adjudication (selection only), `--no-skills` runs with an empty skill library, `--skill ` swaps in another library, `--skills-mode auto` lets the verifier choose which skills to read (see "Skills"), and `--env none` runs without the jail (debugging only). **Run a benchmark.** A cell is one `bench:pool` pair; the runner copies each task workspace under `//_/` and drives the tasks through the driver with per-lane and per-benchmark concurrency caps. It is resumable: finished tasks are skipped, unfinished ones re-staged. ```bash python3 -m harness.runner --run-name demo --cells sb2:opus apex:opus python3 -m harness.score runs/demo/sb2_opus --workers 24 ``` `score.py` reports the single-rollout mean, the selection score, the final score after revision, and the selection oracle. `--select-only` uses only the archived scores and needs no grader; otherwise the delivered bundle and the unrevised base are re-graded in the same pass and `final = select + (revised − base_regraded)`, so grader drift and judge noise cancel in the difference. ## Benchmarks and grading The paper evaluates on APEX-Agents, Workspace-Bench Lite, WorkBuddy Bench, SpreadsheetBench 2 and JobBench. This repository does not redistribute any benchmark data, rollouts or images. `harness/scripts/setup_benchmarks.sh` installs what re-grading needs under `$VERIHARNESS_BENCH_ROOT/benchmarks/`: each benchmark's own code cloned at the commit the paper's graders were validated against, its data from Hugging Face, and its environment and Docker images built with its own tooling. ```bash export VERIHARNESS_BENCH_ROOT=/path/to/benchmarks-root harness/scripts/setup_benchmarks.sh sb2 jb # or: all; --no-images skips the Docker builds ``` | Bench | Upstream | Grader and what it needs at scoring time | | -------- | ------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `apex` | [archipelago](https://github.com/Mercor-Intelligence/archipelago) grading runner | the benchmark's rubric judge with our settings in`harness/benchmarks/apex/`; `GEMINI_API_KEY` (several keys may be comma-separated and are rotated) or `APEX_JUDGE_MODEL=vertex_ai/` with `VERTEXAI_PROJECT` and `VERTEXAI_LOCATION` | | `wsb` | [Workspace-Bench](https://github.com/OpenDataBox/Workspace-Bench) `a3a2ee4` | the benchmark's agent-as-a-judge in its Docker image; judge endpoint in the checkout's`evaluation/.env` (template: `harness/benchmarks/wsb/env.example`) | | `wb` | [workbuddy-bench](https://github.com/Tencent/workbuddy-bench) `b516950` | the benchmark's composite verifier in each task's own image (built per task by the setup script); judge endpoint in the checkout's`.env` (template: `harness/benchmarks/wb/env.example`) | | `sb2` | [SpreadsheetBench-2](https://github.com/RUCKBReasoning/SpreadsheetBench-2) `5c16026` | the official recalculation and cell comparison; LibreOffice in the official image plus`tqdm` (`veriharness-sb2`). The Visualization category is not wrapped | | `jb` | [job-bench-eval](https://github.com/Job-Bench/job-bench-eval) `0152c80` | the benchmark's text judge through an OpenAI-compatible endpoint:`JB_JUDGE_MODEL`, `JB_JUDGE_API_BASE`, `JB_JUDGE_API_KEY` | The three LLM judges (`wb`, `wsb`, `jb`) need an OpenAI-compatible endpoint that you provide: the `.env` templates default to `http://127.0.0.1:4100` and `jb` to `JB_JUDGE_API_BASE`; the proxy in `harness/scripts/litellm.yaml` serves only the verifier lane and can be extended with the judge models. `harness/grade/.py` documents, per benchmark, which upstream entry point is wrapped. Scoring refuses to start when a judge is unreachable (`preflight`), because the upstream judges report a failed call as a 0.0 score rather than an error, and it warns when re-grading the archived bases does not reproduce their archived scores. The `harness/materialize/` adapters build task workspaces from the archived rollout pools of the paper; they document how those pools were produced and are not needed to verify your own rollouts (see "Task workspace" above). ## Data The rollout pool of the paper, about 26,000 rollouts of Gemini 3.5 Flash and Claude Opus 4.8 on the five benchmarks (ten per task), is released on Hugging Face: [caiqizh/veriharness](https://huggingface.co/datasets/caiqizh/veriharness). It contains the rendered trajectory the verifier saw and the files each rollout delivered, with each benchmark's own grader score and a key that maps to the official task id; no benchmark inputs, rubrics or answer keys are included, so the task workspaces are rebuilt from the upstream benchmarks (see the dataset card). `python3 -m harness.materialize` is the adapter that produced the pool from our run archives; to verify the released rollouts, place them in the task-workspace layout described under [Usage](#usage). ## Isolation Verifier sessions run inside `harness/scripts/jail_run.sh`, a mount-namespace jail (`unshare -r -m -p`, no privileges needed): the task workspace is the only visible project state, `spec/`, `workspace/` and `rollouts/` are read-only, the rest of `$HOME`, the data root, `/tmp` and `/var/tmp` are hidden, the Python installation is read-only, and container runtimes are masked. Archived scores and benchmark answer keys are therefore unreachable from a session. There is no network namespace, so model endpoints stay reachable. Graders run on the host after the verifier batch, never inside the jail. The four settings of `--env`: `jail` (default), `none` (no isolation: for hosts without unprivileged user namespaces, or when the data root holds nothing a session should not see; the driver refuses to start in `jail` mode on such a host and says so), `native` and `native-full` (next section). ## Native environments (optional) By default every turn runs in the host jail. With `--env native` the adjudication and delivery turns run inside a fresh container of the benchmark's own task image, so the verifier revises the artifact with the interpreter, libraries and tools the rollouts were produced with; the two investigations stay in the jail. `--env native-full` runs the investigations in the image as well. A task without an image falls back to the jail for every turn. The container sees the task directory at its real path (evidence read-only), the agent runtime and skills read-only, and the host network, which is the same exposure the jail has. ```bash python3 -m harness.runner --run-name demo --cells jb:opus --driver-arg=--env --driver-arg=native ``` Images are resolved per benchmark (`harness/env/__init__.py`); each can be overridden with `VERIHARNESS_IMAGE_`: | Bench | Default image | | --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `wb` | each task's own exported environment image, or its`vh/` derivative when built with `python3 -m harness.env.derive` (the task image plus the harness tool stack) | | `sb2`, `jb` | `veriharness-office`: `docker build -f harness/env/Dockerfile.office -t veriharness-office .` (on top of the SpreadsheetBench image) | | `wsb` | `workspace-bench:local` (the benchmark's image) | | `apex` | `veriharness-docs`: `docker build -f harness/env/Dockerfile.apex -t veriharness-docs .` | Two things a native setup must provide. Task images carry no JavaScript runtime, so the agent runs on a self-contained Node.js 22 placed under `harness/vendor/node-v22/` (the official Linux x64 tarball, 22.19 or later: `curl -L https://nodejs.org/dist/v22.19.0/node-v22.19.0-linux-x64.tar.xz | tar -xJ -C harness/vendor && mv harness/vendor/node-v22.19.0-linux-x64 harness/vendor/node-v22`). And a task image must carry the libraries the skills use, or the verifier's revisions degrade: keep the tool stack in the image, not only the task's own dependencies (`harness/env/derive.py` adds it to WorkBuddy's task images). ## Operating notes * Concurrency: a driver holds up to two model sessions (the investigations run concurrently). Lane limits are a shared budget: a second runner on the same host must be given the same `--max-` values, since each counts the other's drivers. * Grading containers: `wb` and `wsb` grade in networked containers; churning more than about twenty short-lived ones at once destabilises the host's network stack. Use `--batch N` where the grader supports it (`wsb`, `sb2`): a batch of N bundles is judged in one long-lived container, so `--workers 8 --batch 12` grades 96 bundles at a time with eight containers. `wb` grades one task per container; keep it to about 12 workers. * Native mode needs Docker and the benchmarks' task images (WorkBuddy) or the images built from `harness/env/Dockerfile.*`; without them the driver keeps the jail and says so in `driver.log`. Long-lived task containers scale to hundreds without trouble. * Budgets: `--turn-timeout` (1800 s) caps one model turn, `--task-timeout` (3600 s) the whole task; the flash lane on spreadsheet tasks regularly uses the full turn budget in an investigation. ## Skills A skill is a short text (optionally with scripts) describing a reusable failure mode and how to check for it; skills never contain task-specific answers. They are organised by phase: `evidence-*` (tools for reading and comparing artifacts; every phase), `resolve-*` (what evidence settles a disagreement; resolver), `falsify-*` (how a consensus can be wrong; challenger) and `repair-*` (how to revise an artifact without breaking it; delivery). Adjudication receives only the `evidence-*` skills. `--skills-mode` (driver and runner) selects how skills reach a turn: `mounted` (default) gives the verifier the skills that apply to the delivered files; `auto` gives only a catalogue and lets the verifier decide which to read. The runner records the mode in each cell's `run.json`. Provenance: the `evidence-*`, `falsify-*` and `repair-*` texts are the human-authored library of the paper. `resolve-*`, `evidence-patch`, `evidence-bundle`, `repair-patch` and the `xlsx_gaps` / `xlsx_forks` scanners were written afterwards from failure analyses of the harness on the paper's benchmarks, including lessons from the self-evolution experiments on SpreadsheetBench 2 (only reference-free, benchmark-agnostic procedures were kept; scripts that filled cells automatically or encoded a benchmark's reference conventions were not). Those later additions were written from failures on the same benchmarks the harness is evaluated on: no answers or reference conventions entered the library, but the failure modes did, so they are not a held-out test of the library. Each rule is written conditionally on what the task's own text asks for. ## Citation ```bibtex @article{veriharness2026, title = {VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks}, author = {Zhang, Caiqi and Han, Rujun and Wang, Zifeng and CuiZhu, Zoey and Collier, Nigel and Pfister, Tomas and Lee, Chen-Yu}, journal = {arXiv preprint arXiv:2610.00972}, year = {2026}, url = {https://arxiv.org/abs/2610.00972} } ``` ## Contributing See [`CONTRIBUTING.md`](CONTRIBUTING.md). This project follows [Google's Open Source Community Guidelines](https://opensource.google/conduct/); see [`CODE_OF_CONDUCT.md`](CODE_OF_CONDUCT.md). ## License Apache 2.0; see [`LICENSE`](LICENSE). ## Acknowledgements The verifier runs on the [pi](https://github.com/earendil-works/pi) coding-agent runtime (MIT); Claude is served through a local [LiteLLM](https://github.com/BerriAI/litellm) proxy (MIT). Neither is vendored: the setup scripts install both. Re-grading uses the benchmarks' own graders: the [APEX-Agents](https://github.com/Mercor-Intelligence/archipelago) grading runner, [Workspace-Bench](https://github.com/OpenDataBox/Workspace-Bench), [WorkBuddy Bench](https://github.com/Tencent/workbuddy-bench), [SpreadsheetBench 2](https://github.com/RUCKBReasoning/SpreadsheetBench-2) and [JobBench](https://github.com/Job-Bench/job-bench-eval). ## Disclaimer This is not an officially supported Google product. This project is not eligible for the [Google Open Source Software Vulnerability Rewards Program](https://bughunters.google.com/open-source-security).