--- name: eval-report description: 'Write the summary a person reads after an evaluation run: what ran, what it scored, which failures were the MODEL and which were the EVAL itself (out of turns, contaminated, unestablished golden, unreadable judge), what was learned, and what to do next. Every claim links to the artifact behind it. Use at the end of any run, before quoting a number to anyone. Never scores an answer (eval-answer) or decides who owns a failure (eval-diagnose).' --- # Report a Run A run directory is JSONL and a console summary that scrolls away. Neither is a result. This skill turns one into something a person can open, and it is the last step of every run, including a run that failed. The complaint this exists to answer, in a reviewer's words: *"eval runs and a bunch of stuff happens and it is hard to know the actual results."* **Scope boundary:** this skill reports what the ledger already says. It does not score an answer (`skill:eval-answer`), decide who owns a failure (`skill:eval-diagnose`), or edit a model (`skill:eval-improve`). If a number is not in the ledger, do not put it in the report. **The report is about THIS RUN, not about the harness.** Defects you find in the eval tooling while running it are real and worth filing, and they do not belong here: the reader wants to know what their model scored and why, not what is wrong with the thing that measured it. File those against the harness. The one exception is anything that qualifies THIS run's number -- a truncated attempt, a contaminated one, an unestablished answer key -- which the "Eval failures" section exists for. **Do not audit the harness. You were asked to measure a model.** This is the most common way this job goes wrong, and it does not look like going wrong: a run turns up something odd in a script, the odd thing is genuinely a bug, and the reply comes back as a critique of the tooling with the model's score somewhere underneath. The reader asked what their model scored. Answer that. So, unless the user asked you to work on the harness: - Do not read harness source to satisfy your own curiosity about a number. Read it when a number you must report cannot be explained any other way, and stop when it can. - Do not propose harness fixes, refactors, flags or "while I was in there" improvements. Not in the report, not in the chat reply. - When a harness defect DID change this run's number, the report gets one sentence: what the number should be and why. Not the mechanism, not the file, not the fix. - Keep a defect that changed nothing out of the report entirely. Mention it once in chat, in a line, and let the user decide whether they want it chased. A harness bug you found and did not chase is not a loose end. It is the job being done. If the user wants it fixed they will say so, and then it is a different task with its own turn. ## Step 1: build the artifacts, before writing a word ```bash python3 skills/eval-loop/scripts/eval.py package --set --label