--- name: scientific-figure description: > Use when the user has scientific data (or a prompt alluding to scientific data) and wants a publication-quality figure made from it. A generator drafts and renders a figure that lands a frozen communication goal; an adversarial critic critiques it hard and grades it 1-5 per axis against a fixed rubric (message, aesthetic, clarity, integrity, and a conditional domain-completeness axis), aggregates to 0-100, and decides pass; the generator revises against the critic's findings until the grade clears a threshold or the budget is hit. Both roles may consult the literature (Semantic Scholar + arXiv) to verify domain content (e.g. a pathway figure's gene set, or a benchmark's reported numbers), and conform to / grade against a named journal's figure spec fetched via web search. Not for writing a paper or analyzing a dataset, and not for editing an existing finished image — this renders a figure from data/brief and iterates on it. compatibility: Requires Python 3.9+ metadata: version: "0.1.0" --- # Scientific Figure Loop The artifact is a **scientific figure** (the rendered image + the `plot.py` that produces it). Each iteration **generates → critiques+grades**: a **generator** authors a rendering script and renders the figure to land the frozen `` message; an adversarial **critic** grades it 0-100 against the fixed `rubrics/rubric.md` and decides `pass`; the generator then revises against the critic's concrete `findings`. The loop runs until the grade clears `` or the budget is hit. All work happens on copies inside a sandbox; the user's data is copied in read-only and never edited. The cast (all in this folder): - `roles/generator.md` — drafts/revises `plot.py`, renders `figure.png` by running ``, optionally grounds domain content via ``; writes `generation_notes.md`. - `roles/critic.md` — the adversarial grader: re-derives each rubric axis independently, spot-checks the figure's numbers against the data, optionally lit-checks domain completeness, and emits `schemas/critique.schema.json` (the grade + `pass` + executable findings). - `rubrics/rubric.md` — the **fixed** grading rubric (the critic never edits it). - `schemas/critique.schema.json` — the one validated output. **Spawn-or-degrade.** On Claude Code, spawn the `generator` then the `critic` as real `Agent` subagents (sequential — the critic needs the generator's figure); otherwise adopt each role inline. You are the orchestrator. ## Why the critic grades itself (the honesty problem) The critic both critiques **and** grades, which under loop-termination pressure invites inflation and a generator that games the rubric. `roles/critic.md` + `rubrics/rubric.md` counter this: the critic (1) applies a **fixed** rubric it never edits, (2) **re-derives** each axis from the rendered figure + data + frozen `` rather than echoing the generator, (3) **recomputes** a sample of the figure's numbers itself instead of trusting "it's fixed", (4) holds a **fixed, anchored** bar with **no credit for effort or elapsed iterations**, and (5) applies **hard gates** (a figure value that contradicts the data, a misleading axis, or fabricated data presented as real fails the figure regardless of the average). The generator optimizes the concrete `findings`; the critic grades holistically against the frozen goal — so "address every finding" does not mechanically buy a pass. Because the two are separate agents, the critic never just rubber-stamps the generator's intent. ## When to use Use when scientific data (or a prompt describing it) exists and the user wants a polished figure pushed past a quality bar with adversarial critique and a graded rubric. Default: run the full generate→critique loop below. Escape hatch: if the user only wants one figure + a critique (no iterating), run one generate + critic pass and stop. Not for writing a paper or doing the analysis, and not for retouching an already-final image. ## Setup **Resolve bindings interactively.** If `loop.run.yaml` exists in the working dir, load it, confirm the values in one line, and skip to the loop. Otherwise: on Claude Code (the `AskUserQuestion` tool is available) infer a likely value for each binding and present it as the recommended option; on other hosts ask each as a quoted plain-text prompt. Then write `loop.run.yaml` (format: `examples/run.example.yaml`) and confirm every value — including the distilled ``, whether a journal spec applies, and the live/degraded literature tier — before creating any other files. | binding | meaning | default | how to infer | |---|---|---|---| | `` | the prompt describing the figure to create + the message/claim it must communicate (and, if data exists, what the data represents) | — | the user's request; if pasted as prose, save to `/brief.md` | | `` | data file(s) the figure visualizes (CSV/TSV/parquet/JSON…); **empty → an illustrative/schematic figure** (the integrity axis then checks internal consistency, not data fidelity) | — | scan the working dir near the request; may be null | | `` | the figure's communication objective(s), 1-3 bullets — **frozen**; the critic grades against these and the generator may never abandon them | — | **distill from `` at setup**, confirm with the user in one line | | `` | command/interpreter that runs the plot script the generator writes (it appends `iter/plot.py`), in the user's env — the skill ships no plotting deps, the same contract as scientific-writer's `` | `python3` | `pyproject.toml`/`.venv`/README; e.g. `uv run python` or a venv python | | `