--- name: eval-loop description: 'Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint. You are the conductor: import cases into the file ledger, spawn a blind answerer, then run eval-answer, eval-diagnose, and eval-improve. Persistence is plain files: the set in the model package''s evals/ directory, runs in the set''s workdir; checkpoints are git commits of the model repo. `eval.py` runs each step from the set''s eval.toml. Use to score a model, diagnose failures, improve behind an acceptance check, or roll back a bad direction.' --- # The Evaluation Loop You conduct this loop. There is no batch orchestrator to start, no eval API, and no eval MCP tools. The ledger is plain files: the set in the model package's git repository, and each run in the set's workdir (`reference/ledger-schema.md` in `skill:eval-answer` defines every file and event). `scripts/eval.py` runs each step from the set's `eval.toml`; `reference/running-a-run.md` has the commands. Scoring is an LLM judge you spawn per case. There is no scripted scorer, and there will not be one: a script that can pass a wrong answer is worse than none. The scripts under `scripts/` run the loop -- they answer, re-execute, spawn the judge, compare runs, and write the ledger -- but none of them decides whether an answer was right. ``` scrape/run -> eval -> diagnose -> improve -> checkpoint ``` **This skill conducts; it does not restate.** Scoring lives in `skill:eval-answer`. Components and owners live in `skill:eval-diagnose`. Edit rules live in `skill:eval-improve`. Do not merge **eval** into **diagnose**. A conductor who scores while explaining writes the explanation into the score. Do not skip the **acceptance check** inside improve. The acceptance check decides whether *this* edit stays. **Checkpoint** decides whether a *sequence* of accepted edits can be undone. ## Where the rest of this lives This file is the procedure. The things it used to carry inline are files beside it now, because each is needed at one moment rather than every run, and loading all of them for every run is how a skill stops being read. | When | Read | |---|---| | the set does not exist yet, or has never run | `reference/setting-up-a-set.md` | | about to run one | `reference/running-a-run.md` | | a golden is wrong, doubted, or out of step with the model | `reference/golden-side-door.md` | | auditing a key you doubt, or a set you did not author | `reference/auditing-an-answer-key.md` | | deciding whether an edit stays | `reference/acceptance-check.md` | | about to quote a number, set the band, or read a set's flips | `reference/measurement.md` | | the run finished and someone has to read it | `skill:eval-report` | | you changed judge doctrine or its inputs | `reference/checking-the-judge.md` | Read the file, do not work from the summary here. The acceptance-check rules and the golden side door are both places where acting on a half-memory of the rule produces a confident wrong answer rather than an error. ## The five steps | Step | Job | Writes | |---|---|---| | **a. scrape / run** | Put cases in the ledger; spawn a blind answerer | cases; `attempt`, `tool_call` | | **b. eval** | Judge the answer; score which required entities retrieval delivered | `score` | | **c. diagnose** | Why it failed, who owns it | `issue` / `issue_status`. Stop. Do not edit. | | **d. improve** | One smallest model edit, then the acceptance check | improve writes `candidate`; you write `acceptance_check`. Revert on reject. | | **e. checkpoint** | Git commit after an accepted acceptance check | `checkpoint` event, then the commit | Every run then **reports** (below). A run whose result exists only as JSONL and a scrolled-away console summary has not been delivered. **scrape** and **run** share a letter but are not the same job. Scrape writes cases. Run writes attempts. Do not invent questions and score them in one breath. **Step d is `skill:eval-improve`, not a hand edit plus another arm.** The failure mode is specific and it has happened: the edits were made by hand, published, and validated by re-running the whole arm -- with the answer key repaired in the same window. No `candidate`, `acceptance_check` or `checkpoint` event exists in that run's ledger, and the resulting move cannot be attributed to the model, the key, or the answerer. Re-running a whole arm after changing two things is the one thing the acceptance check exists to replace: it scores the edit against dev AND holdout, and it is cheaper than the arm. ### Scrape, minimally Importing an existing corpus IS the scrape step: copy the set from its home (for example a benchmarks checkout) into `evals//` and convert to the ledger shapes. **`skill:eval-import` is that job in full** -- how to classify what arrived with each question, why nothing imports as a verified value, and the seal that makes a later edit to a question detectable. Read it whenever the questions came from outside, which is most of the time. While importing: - Freeze each case's `split`: `dev` or `holdout`. Diagnose and improve read dev cases only; the acceptance check runs both. A set that is all dev cannot defend an accept. - Later, each diagnosed-and-fixed failure becomes a new frozen dev case, so a fixed bug cannot silently return. Scraping from production logs (chat transcripts, retrieval traces) is the other supported source, and usually the better one: real traffic asks what people actually ask. Where your logs physically live is a host concern; look for a host-specific log-fetching skill. `skill:eval-import` takes over once you have the text, and its `reference/case-format.md` covers what a log pull needs that a question list does not. Prefer variety over volume when you sample, from either source. Cases that differ in grain, source, filter shape, and phrasing are what move a measurement; a second sample of the same case is nearly free of new information. ### Mode aliases Older mode names still work as aliases for how far one run walks: | Alias | Steps | |---|---| | `measure` | scrape/run + eval (`eval.py package` then needs `--without-diagnosis`) | | `triage` | plus diagnose | | `improve` | plus improve + acceptance check + checkpoint on accept | Say which alias (or which steps) you are running before the first question. Record it in `run.json`. Do not mix steps in a way that lets the answerer see gold, issues, or the model file. Most runs should stop after eval. Diagnose when you need a histogram of components and owners. Improve only for diagnosed *model* gaps, one batch at a time. Checkpoint only after the acceptance check **accepts**. ## Roles | Role | Sees | |---|---| | **Answerer** | The question and the Malloy tools. Never the golden, `evals/`, the model file, or any hint it is being evaluated. | | **Judge** | The golden and the prediction. Never conducts, never answers, never edits. One fresh subagent per verdict (`skill:eval-judge`). | | **You (conductor / improver)** | Everything, including goldens and traces. | | **Acceptance check** | The edit and the evidence. Never the improver's self-assessment alone. | The answerer stays blind. That is not optional. A grader-visible answerer writes toward the expected answer, and the score is fiction. There are no eval MCP tools on purpose. The answerer inherits your tools, including Shell and Read, so any eval convenience surface would also be a gold path for it. Blindness is prevention plus detection, not a guarantee: `eval-answer` runs the contamination checklist on every attempt, which is why you keep a host-side tool-use log per answerer. ## Pick the target first Both a local model server and a hosted platform expose the same two tools the answerer needs, `get_context` and `execute_query`, so the loop runs against either. What differs is which model is answering and whose data it reads, and those are two separate axes: | Target | Model under test | Data | Can edit and re-test? | |---|---|---|---| | **Local (direct)** | your working files | local (for example duckdb), or a direct warehouse connection | yes | | **Local (proxied)** | your working files | the platform's connection, through a proxy connection type | yes | | **Remote** | the published version, through the platform's hosted `get_context`/`execute_query` | the platform's | no, publishing is not an eval action | The middle row is the one worth knowing about: it decouples the two axes, so you can evaluate a model you are still editing against the customer's real data. It is a connection configuration, not a feature. ### Three ways to reach the tools That table is about which MODEL answers. A second, independent choice is where the answerer's MCP tools come from, and `--mcp-url` is the whole of it: | | `--target` | `--mcp-url` | Auth | |---|---|---|---| | **1. Local Publisher** | `local` | `http://localhost:4040/mcp` (default) | none | | **2. Hosted, through an editor extension's local bridge** | `platform` | the localhost URL the extension prints | none: the extension holds the credential | | **3. Hosted, directly** | `platform` | the host's `https` endpoint, scoped if it offers one | a cached OAuth login, once, interactively | ```bash # 1. local --target local # --mcp-url defaults to the local Publisher # 2. hosted via the extension's bridge -- no OAuth, but check what it exposes: # the same proxy may front a local Publisher instead --target platform --mcp-url http://localhost:/mcp \ --hosted-mcp-server --target-version --scope /@ # 3. hosted directly -- authenticate first, under the SAME server name claude mcp add --transport http claude # /mcp -> -> Authenticate --target platform --mcp-url \ --hosted-mcp-server --target-version --scope /@ ``` All three hand the answerer the same three capabilities (`get_context`, `execute_query`, and the docs search), so a comparison between them is between agents that could do the same things; `test_the_two_arms_hold_the_same_capabilities` pins it. Modes 2 and 3 take those names from `--hosted-tools`, which defaults to the bare trio; pass it only if this host names them differently. A platform run left on the local default `--mcp-url` is refused rather than probed, because pointing every answerer at a local Publisher measures a different model over different data than the run claims. Two rules follow, and both are the kind of mistake that produces confident nonsense rather than an error: - **The answerer and the conductor must hit the same target.** If the answerer queries the published model and you re-execute its query against your edited local copy, the score describes neither. Decide the target before the first question and record it. - **Pin the version the target actually served, not the one you happen to have.** A local target pins a commit; a platform target pins the published version. Recording a local commit for a run that queried a published model is a pin that means nothing. Which target for which job: - **Baseline what customers experience:** Remote. It is the deployed model through the deployed engine, which is the thing they actually hit. The judge sees no re-executed rows on a Remote run (there is no local copy of the bytes), so its verdicts rest on the answer text and the golden; say so. - **Improve and accept:** local, because the acceptance check needs compile, reload, and a fresh re-answer between edits. Publishing to a customer environment to score an edit is not something this loop does. Where the host offers draft execution, that counts as local for this purpose. - **Measure real data without touching production:** local proxied. So a measure-only run can use any target; a run that includes **improve** needs a local one. Two things to check before a platform run, because neither errors and both make the run measure something other than what it names: - **The answerer's skills must be written for THIS host.** A shared skill names an MCP tool by its bare name (`get_context`) so it reads correctly anywhere, but a host/router skill names its own host's tools directly. Install the latter for the wrong host and the answerer is told to call tools it does not have. `run_baseline.py` warns when the manifest it loaded names Publisher-only tools on a platform target; point `--answerer-manifest`, or `--skills-root`, at the checkout that ships this host's manifest. - **The tool names are configuration.** `--hosted-mcp-server` is both the `mcp____` prefix and the OAuth cache key, so it has to match the name the answerer authenticated under, and `--hosted-tools` lists the bare tools that host exposes. - **Get the hosted tools in front of a headless answerer, one of two ways.** A spawned answerer cannot complete an OAuth flow, so the tools have to be reachable before the run starts. `run_baseline.py` proves it with one cheap probe and refuses to spend an arm otherwise -- a run whose answerers have no tools does not error, it reads as a terrible model. 1. **Authenticate once, interactively.** Works anywhere, including a plain CLI install, and is the route to assume unless you know otherwise. The token is cached per server NAME, so authenticate under the same name the run passes to `--hosted-mcp-server`: ```bash claude mcp add --transport http claude # then /mcp -> -> Authenticate ``` Then come back and run. This is a hand-off to a person; there is no headless equivalent, so plan for it rather than discovering it mid-run. 2. **A local proxy that already holds the credential.** Some hosts ship an editor extension whose local MCP proxy can expose the hosted `get_context` / `execute_query` -- often behind a setting that is off by default. Where that exists, point `--mcp-url` at the proxy on localhost and no OAuth step is needed, because the extension holds it. Check what the proxy actually exposes before relying on it: the same proxy may serve a local Publisher's tools instead, and then `--hosted-tools` is naming tools that are not there. This route is not available to someone running the CLI alone. - **Prefer a SCOPED endpoint URL over asking for scope.** A hosted MCP is usually reachable two ways: a global endpoint where every call carries an organization and workspace, and a scoped one where the URL itself is the scope. `--scope` and the prompt can only ASK an answerer to stay in one package; a scoped URL enforces it. For an agent being measured that is the difference between a case answered against the package it names and one answered against whatever else the account can see. Authenticate once interactively (`claude`, `/mcp`) under the same server name the run will use; the token is cached per name, and a spawned headless answerer cannot complete an OAuth flow. - **Pin the VERSION in the scope, not just the package.** Write `--scope /@`. Both hosted tools take a version and both document the same default for an omitted one: the PINNED version, which is whatever the workspace serves at the moment of the call. So an unversioned run records `targetVersion` in `run.json` and then answers from whatever is current, and the two part company the moment anyone publishes -- including mid-run, which measures two builds under one label. `--target-version` fills the version in when the scope omits it, so a platform run is pinned without opting in; a scope naming a different version is refused rather than taken as an override. ## Before you start 1. The model package under evaluation must live in a git repository, with `evals//` in the package, beside the model files. Git is the checkpoint mechanism; without it there is no rollback and no run can include improve. Keeping the set IN the package is what stops a model edit and its answer key drifting apart: they move in one commit, so fixing a measure and forgetting the golden that depended on it stops being possible. It is safe -- measured on a running server, a `cases.jsonl` inside a package appears in no model listing, no notebook listing, no package resource, and 404s over HTTP, so an MCP-only answerer has no route to it. What it buys differs by target. On a LOCAL Publisher it does not get you free versioning -- `sourceContentSha` hashes model paths only, so the set needs its own `datasetSha`. On a hosted target that publishes the whole package directory as an IMMUTABLE version, the set rides inside that version and `targetVersion` pins model and answer key together; nothing can be edited under a published version, which is what makes it a pin. Check which you have before deciding how much of this you need. **Look for a set and for prior runs before you author either.** A minute of `find . -name cases.jsonl`, a glance at the set's workdir (`/runs/`, `~/.malloy-eval/-/runs/` unless `eval.toml` moves it) and at your host's own transcripts for this repo. Two sessions fourteen minutes apart built the same 29-case answer key from scratch, because the first had committed nothing before it was deleted and the second had no way to know it existed. Roughly a working day was spent twice, and three specific things were rediscovered at cost: the MCP login flow, the retrieval gate's 401, and a judge-rendering bug that had already cost $17 of arm once. The two keys, independently authored from the same questions, differ by about ten points on comparable answers -- which is the available measure of how much a key depends on its author, and a reason to reuse one rather than rebuild it. **Audit the set's entity ids before the first arm.** `verify_goldens.py` checks that each id names something in the model; `check_findable.py` checks that a search of its own kind actually returns it. Both are free of model calls. An id that fails either scores a retrieval miss on every run, and the miss reads as the model's fault. **Commit the set before you spend money on an arm**, and keep durable outputs in the repository. A findings document in `~/Downloads` is gone the first time somebody tidies up; the set and the write-up belong in git beside the model. **Runs go in the set's workdir, never inside the model package.** A run directory holds a `model.malloy` snapshot, and a built report is a Malloy package; inside the package under test, either puts that package into `loadErrors`. The default workdir, `~/.malloy-eval/-/`, is outside git, which suits a measure-only run. A run that will improve and checkpoint needs its ledger kept: set `[paths] workdir` in `eval.toml` to a directory in the model's repository but outside the package, and gitignore its `servers/` and `packages/`, which are a server's database and rebuilt reports. 2. The server must be up. Where your host offers retrieval tracing, turn it on and confirm a trace lookup answers, so a call's ranked results can be recovered afterwards. **Open-source Publisher has no trace lookup.** `PUBLISHER_MCP_TRACE=retrieval` is set by `serve.py` and makes Publisher write one "Retrieval trace" line per ranked call to its log (stage counts and timings, not the ranked results). Nothing reads that line yet: there is no trace store and no trace tool, so `traceId` is null on every local attempt. This does not block a scored run, because attribution never depended on it: `rankedSummary` is copied onto the `tool_call` event at capture, precisely so the evidence survives without a store. So do not refuse a local run for want of tracing -- an earlier version of this rule did, and it refused every local run there has ever been. Refuse one whose `tool_call` events carry no `rankedSummary`, which is what attribution actually reads. 3. Health-check: your host's status check until it reports serving, and inspect `loadErrors`. A dead database that still answers HTTP is an environment failure, not a model failure. Stop and fix it. Four consecutive environment or no-result attempts means stop the run. 4. Load the set: scrape/import as above, or reuse an existing `evals//`. `eval.py check --set ` names every gap in it before anything starts. Never keep two live copies of one set; the set directory in the model repo is the single source of truth, versioned by `datasetVersion` in `set.json`. 5. Review goldens before you score. **Check how many cases can take a verdict at all**, not just how many cases there are: a golden the set stamps `provisional`, `invalid` or `ambiguous`, or a case with no golden, scores `verdict: null` and stays out of the pass rate. An imported set is `provisional` throughout by design (`skill:eval-import`), and the only thing that changes that is `verify_goldens.py --promote` after a re-derivation through the truth package. A run whose every key is underived is refused rather than spent; a set of bare questions runs, because its answers are what keys get derived from. A verified golden that holds a value but no local artifact stays verified by provenance and is not scorable until you have rows or a scalar to compare (the judge needs both sides). A `criteria` golden is the exception: it holds no value, so its clauses are both sides. If diagnosis later marks `BAD-REFERENCE` or `AMBIGUOUS-REFERENCE`, follow `reference/golden-side-door.md`. Both are expected in the wild; both are the golden side door below, not improve, and not a sixth step. 6. `eval.py run` creates `/runs//run.json` with the attribution pins (`reference/ledger-schema.md`): mode, dataset version, **the target and the version it served** (a local target pins a commit, so commit or stash first; answering from a dirty tree pins nothing), server version, judge version and rubric sha, answerer model, call budget, trace mode. Freeze those for the whole run. Raising a call budget mid-run moved mean outcomes on an unchanged model. The call budget is `--max-turns`, written to `run.json` as `maxTurns`. **Size it from a pilot rather than taking the default of 30.** Run the three cheapest cases uncapped (`--max-turns 100`) and set the cap at twice their maximum. A set whose questions need two or three sources joined does not fit a cap sized for single-source lookups: on one such set the completed attempts had a median of 17 turns and a 90th percentile of 26 against a cap of 30, and four cases died at it. A cap 15% above the 90th percentile of completed work is not a safety margin. Note also that `malloy-analysis` tells the answerer to persist -- retry a phrasing, let a small query settle whether a field exists -- so a tight cap and that instruction are in direct conflict. **Do not change the model and the measuring instrument in the same step.** Those are two different axes and only one of them is cheap to separate. Batching MODEL edits is fine and expected. Clustering exists so that one edit closes several cases, and an arm per fix does not survive contact with arithmetic: 20 fixes over 100 cases is 2,000 answers, and at the measured $0.33 a case on a proxied warehouse that is $660 of answering to attribute what the acceptance check attributes for the price of the affected cases plus holdout. Budget five arms for a defensible claim (a baseline, two for the A/A, two post-edit), not one per edit, and let `skill:eval-improve` carry each cluster. The instrument is the other axis: the answer key, the judge and its prompt, the answerer model, the skills. Move one of those together with the model and there is nothing left holding still, so the result measures neither. On the run this comes from, nine model commits and a re-derived answer key landed between two arms, and the move from 20% to 39% belongs to no one -- not because two model edits were batched, but because the ruler changed at the same time as the thing being measured. A key repair mid-improve is not forbidden; it ends that comparison, so re-baseline rather than quoting a delta across it. The answerer model is instrument too, and it is the one most often left unstated. The same 23-case set read 12 match / 11 near / 0 no_match on one model and 8 / 4 / 15 on a smaller one. No statement about "the agent's" capability means anything until two arms name the same answerer. 7. Generate every answerer prompt from the stored case in `cases.jsonl`. Never retype the question. A truncated retype is indistinguishable from a real question downstream. ## Per question 1. Health-check again. 2. Spawn a *fresh* blind subagent. Give it only the question text and the Malloy analysis tools. Tell it to follow the `malloy-analysis` skill. Do not mention eval, gold, scoring, or this skill. 3. Keep a host-side tool-use log for that subagent (name, input path or command, MCP tool name). Publisher traces see MCP only; a Read of a gold CSV is invisible server-side. 4. `skill:eval-answer`: contamination first, then re-execute, then the judge, then events. 5. `skill:eval-diagnose` only when this run includes diagnose, only on dev cases, and only after the score event exists. 6. `skill:eval-improve` only when this run includes improve, and only for `owner: model`. Then run the acceptance check. On accept, checkpoint. ## Report the run, or nobody can read it A run directory is JSONL. It is a record, not a result, and the console summary scrolls away. **Every run ends by producing something a person can open**, and that job is `skill:eval-report`: it builds the servable run package (the case matrix app and the aggregate notebook) and gives the template for the write-up. Read it at the end of every run, including a run that failed. The standing complaint about this loop is that "a bunch of stuff happens and it is hard to know the actual results", and a ledger nobody renders is why. Two rules from it are worth repeating here, because they are the ones a conductor skips: - **Separate a MODEL failure from an EVAL failure.** A wrong answer and a broken measurement look identical in a pass rate and have nothing else in common. An arm holding a truncated, contaminated or environment-failed attempt has no rate to quote at all. - **Say when a step did not run, and why.** `improve` not running because every cluster came back `owner: agent-skill` is a RESULT, and it reads identically to having forgotten unless it is written down. Keep the write-up in the repository beside the set, not in a chat log and not in `~/Downloads`. ## Checkpoint A checkpoint is a git commit of the model repository, taken after an acceptance check accepts, so a bad improve direction can be rolled back. It is not a report, and it is not a remote publish. 1. Commit the model files AND the set's ledger in one commit; put the label and the closed issue ids in the message. The run's ledger is in that commit only when the workdir is in the repository (Before you start, item 1). 2. Append the `checkpoint` event (`action: created`, label, `modelGitSha` from the commit you just made, issueIds). The event line itself rides in the next commit; append-only logs trail by one commit and that is fine. 3. Confirm `git status` is clean for the model files. **Restore**: `git checkout -- ` (or `git revert` the checkpoint commits), then reload the package, then append a `checkpoint` event with `action: restored` and the sha. Readers return to the model that existed before the bad direction. Take a checkpoint of the current model *before* the first improve batch if no commit pins it yet. Rolling back by hand is guesswork. If reload reports `mode: reinstalled`, the package was re-fetched from its install location and may have overwritten the restored files. Prefer in-place / watch-mounted packages for this loop. ## Out of scope This loop is local. The ledger is files, the checkpoints are git, you are the conductor. Do not: - publish the model to a hosted platform as a "true" checkpoint or learning curve - start a Python orchestrator (`loop.py`, `run_all.py`, `improve_batch.py`) that runs the five steps end to end unattended. You conduct; the scripts are the steps, not the sequencing. There are more than twenty of them and they are not the exception to this: each does one step you invoke and hands back a result you read. What is forbidden is a script that decides what to do next, because every judgement this loop protects lives in that decision - score by string-diffing rows instead of judging them, or reintroduce a scripted row oracle: one that can pass a wrong answer is worse than none - wait for a bigger gold set before the loop can run; dev/holdout on what exists beats waiting - register eval MCP tools or stand up an eval API - encode unsettled goldens into the model ## Prime directives - The model is the only thing improve edits. No question text, qids, or expected values in any name, doc, or comment. - **You are measuring a model, not reviewing this harness.** Read a script when a number you have to report cannot be explained otherwise, and stop there. Do not audit the scripts, propose fixes to them, or hand back tooling critique in place of a result -- a run that ends in harness feedback has not answered the question it was asked. A defect that changed a number gets one sentence in the report; a defect that changed nothing gets one line in chat, once. Fixing it is a different task, and the user starts it. - When the environment misbehaves, stop. Never diagnose a sick system. The harness is part of the environment: a known-broken measurement does not become quotable by being finished. Measured against this directive, an arm already paid for gets quoted anyway -- it happened three times in one run, once with the confound stated in the same message that started the arm -- so the harness now enforces the parts it can. A run with a truncated or contaminated attempt prints no pass rate and records `status: incomplete`, and `flip_table.py` names what one arm left unscored that the other did not. Read `run_error` on the attempts and the INCOMPLETE line before quoting any number. - When a subagent disagrees with you, probe. Do not win by authority. - When a rule here is wrong, change this file and note it on the run. ## Related skills - `skill:eval-answer`: contamination, judge protocol, events. Its `reference/ledger-schema.md` is the file contract; `skill:eval-judge` is the judge. - `skill:eval-diagnose`: component, owner, issue events. No edit. - `skill:eval-improve`: smallest model edit, probe receipts, no self-accept. - `skill:eval-report`: the run package and the write-up a person reads. - The `malloy-analysis` skill: what the blind answerer follows. It is installed from the `analysis` manifest group, not the `eval` group.