--- name: run-model-eval description: Use when running the Eval-v1 agent-reliability benchmark (packages/web/eval) — benchmarking a newly released Ollama Cloud model, re-running or adding scenarios, refreshing the published dataset in packages/web/eval/data/, or when a long sweep needs unattended keep-alive (watchdog), stalls, or produced suspicious/contaminated results. --- # Run Model Eval (Eval-v1, pinchy#669) ## Overview State-based agent-reliability benchmark: real models over Ollama Cloud `/v1` drive a Pinchy agent against mock email/ERP backends; grading reads the database back, never the transcript. Harness mechanics live in `packages/web/eval/README.md`; the published dataset contract in `packages/web/eval/data/README.md`. This skill is the **operational runbook**: the ordering, the iron rules, and the gotchas that are NOT recoverable from the repo alone. Core principle: **probe before you sweep, one sweep per stack, everything resumes from JSONL.** ## Iron rules (each one cost us real damage once) 1. **REFRESH THE CATALOG FIRST.** Before ANY sweep, run `pnpm models:discover` (see the `update-ollama-cloud-models` skill) and act on the delta. The model set decays under you: on 2026-07-15 Ollama retired `deepseek-v3.2` and `glm-4.7` mid-benchmark, and we only noticed two days later — by accident, while researching prices. `models:discover` exits non-zero on `REMOVED`, so this is a 30-second check that prevents two expensive failures: a sweep that burns hours 404-ing on a model that no longer exists, and a published benchmark whose model set the provider no longer serves. `ADDED` matters just as much — a sweep that silently omits the newest models is stale the day it ships. The skill's own trigger list said "before a release", never "before a sweep"; that gap is exactly how this bit us. Retired models are NOT deleted from the dataset: their last measured numbers stay published and citable, marked as withdrawn from the serving path (see `data/CHANGELOG.md`, and the legacy policy in `data/README.md`). 2. **PROBE FIRST.** Before any full sweep of a new model or scenario: run N=3 × 3-4 capable models (`EVAL_N=3`, `EVAL_CANDIDATE_MODELS=...`), then **read the trajectories** (`results/