--- name: goal-test description: >- Turn a vague task into a testable definition of done and generate an executable goal-test script for it, optionally with a bounded retry loop around a headless agent. Use when the user wants to run an agent in a loop, asks how to know when an agent task is finished, or says a goal like "improve X" needs to become checkable. Do NOT use for building eval suites over many cases — that is evals-bootstrap. --- # Write done as a script Theory: [The four parts of a loop](https://secondbrainos.dev/#/read/course-2-loop/the-four-parts) and the [goal test build page](https://secondbrainos.dev/#/read/track-loop/build-goal-test). "Improve the error handling" cannot terminate a loop, because nothing can ever say it is finished. The goal must be phrased so that a program — not a person, not the model — returns true or false against it. ## Workflow 1. **Extract the claim.** Ask what the user wants true at the end, then restate it as verifiable facts: which command exits 0, which file exists, which string appears, which number crosses which threshold. If a part cannot be checked by a program, say so and negotiate it down to what can. 2. **Prefer checks that already exist.** The test suite, the linter, the build, the type checker. A goal test that shells out to `make test` is better than one that reimplements it. 3. **Write `goal-test.sh`** (or `.py` if the checks are easier there) in the project root: runs every check, prints one line per check with pass/fail, exits 0 only when all pass. Keep it under ~40 lines; it must run in seconds and be safe to run repeatedly. 4. **Run it now.** It should fail before the work is done — a goal test that passes on the current state is testing nothing. Show the failing output. 5. **Offer the loop.** If the user wants the agent driven until done, wrap it: ```bash #!/usr/bin/env bash # goal-loop.sh — run a headless agent until the goal test passes or attempts run out. set -u MAX_ATTEMPTS=5 for i in $(seq 1 "$MAX_ATTEMPTS"); do FEEDBACK=$(./goal-test.sh 2>&1) && { echo "done in $i attempt(s)"; exit 0; } claude -p "Goal: . The goal test currently fails with: $FEEDBACK Fix the code so the goal test passes." \ --permission-mode acceptEdits --output-format json > ".attempt-$i.json" done echo "goal test still failing after $MAX_ATTEMPTS attempts — falling back to a human" exit 1 ``` Adjust the agent command to whatever CLI the user runs. Keep the three brakes visible and named: the checker outside the model (`goal-test.sh`), the stop rule (test passes), the budget (`MAX_ATTEMPTS`, plus a dollar cap read from the JSON output if they want one). ## Rules - The checker never asks the model's opinion; self-review is not a checker. - Feed the checker's output back verbatim — the failure text is the prompt. - If the user cannot state done as a check, they do not have a loopable task yet; tell them that plainly and suggest an interactive session instead.