--- name: computational-experiments description: "Scaffold, execute, analyse, and publish computational research experiments through a reproducible staged workflow. Use when a research question requires simulations or computational sweeps rather than a one-off script." allowed-tools: Bash(uv*, pytest*, mkdir*, ls*, cp*), Read, Write, Edit, Glob, Grep, AskUserQuestion, Skill argument-hint: "[project-path] [--mode scaffold|experiment|figures|full] [--budget ] [--scaffold standard|robustness|replication]" agent-dependencies: [code-review] skill-dependencies: [multi-perspective] --- # Computational Experiments > Lifecycle skill for algorithmic research projects where the code IS the scientific contribution. ## Modes | Mode | What it does | Phases | |------|-------------|--------| | **Scaffold** | Create/audit package structure + algorithm skeleton | 1–2 | | **Experiment** | Design and run pre-specified sweep campaigns | 1, 3–4 | | **Explore** | Adaptive experiment loop: modify → run → evaluate → keep/discard | 1, 3E–4 | | **Autonomous** | Parallel self-correcting sweep with sub-agents | 1, 3A–4 | | **Figures** | Generate publication output from results | 1, 4 | | **Full** | Complete pipeline | 1–5 | Default: **Full**. Detect mode from user request or ask if ambiguous. ### `--scaffold` Flag Sets a **stage progression template** for the experiment campaign. Templates provide structured checklists and exit criteria for each stage. | Scaffold | Stages | Best for | |----------|--------|----------| | `standard` | Init → Tune → Creative → Ablate | Algorithm development, ML, simulation | | `robustness` | Main spec → Alternatives → Placebo → Sensitivity | Causal inference, econometrics | | `replication` | Exact → Our data → Extensions → Robustness | Replicate-and-extend papers | Templates live in `templates/experiments/`. When a scaffold is active: 1. Present the current stage's checklist before starting work 2. Gate progression: don't move to the next stage until exit criteria are met 3. Log stage transitions in the experiment breadcrumb Default (no flag): user-defined stages (current behavior). Scaffold is a guide, not a cage — users can skip stages or reorder with explicit acknowledgment. ### `--budget` Flag Sets a **campaign-level time budget** in minutes for the entire experiment run. When set: 1. **Record start time** at the beginning of Phase 3 (any variant) 2. **Check remaining budget** before launching each new config, sweep batch, or explore iteration 3. **Soft stop** when budget is exhausted: finish the current run, save all results collected so far, skip remaining configs 4. **Never hard-kill** a running experiment mid-execution — always let the current run complete | Mode | How budget applies | |------|-------------------| | **Experiment** | Skip remaining sweep configs when time is up | | **Explore** | Exit the explore loop at the next iteration boundary | | **Autonomous** | Do not spawn new agent batches; let running batches finish | | **Figures** | Budget does not apply (figure generation is fast) | **Reporting:** When a budget triggers early stop, append to the breadcrumb: ``` - **Budget:** Stopped after / configs (budget: min, elapsed: min) ``` **Default:** No budget (run all configs to completion). Typical values: `--budget 30` for quick exploration, `--budget 120` for overnight sweeps. Use **Experiment** when the design space is known upfront (grid/random sweep). Use **Explore** when the design space is unknown and the agent should adaptively search (inspired by Karpathy's [autoresearch](https://github.com/karpathy/autoresearch)). ## When to Use - "Set up experiments for my algorithm" / "Scaffold the project" - "Run the benchmark" / "Design the sweep" - "Explore what auction parameters reduce collusion" / "Try different architectures" - "Generate convergence plots" / "Make publication figures from results" - Any project where algorithms, simulations, or optimisation loops are the contribution ## When NOT to Use - Empirical data analysis (observational/survey data) → `data-analysis` - Experimental design for human subjects → `experiment-design` - Causal inference strategy → `causal-design` - LaTeX compilation → `latex` ## Workflow ### Phase 1: Detect & Configure 1. **Read project context:** `CLAUDE.md`, `MEMORY.md`, `.context/project-recap.md` if they exist 2. **Detect structure:** Scan for `src/`, `experiments/`, `tests/`, `pyproject.toml`, `setup.py` 3. **Detect language:** Check existing files. Default: Python with uv. Ask if ambiguous. 4. **Read MEMORY.md:** Look for `[LEARN:code]` tags, notation registry, key decisions — apply established conventions 5. **Determine mode:** From user request or ask. If existing package exists, skip scaffold. 6. **Inventory existing code:** Map modules, entry points, config objects, result files Read `references/package-scaffold.md` if scaffold mode is active. ### Phase 2: Scaffold **Skip if:** Package structure already exists and user requested experiment/figures mode. 1. **Package structure:** Create `src//`, `tests/`, `experiments/configs/`, `scripts/` 2. **Build config:** `pyproject.toml` with hatchling, dev dependencies (pytest, matplotlib, numpy) 3. **Algorithm skeleton:** Base classes (`Algorithm`, `Experiment`, `Metric`) with `# TODO:` markers. For multi-agent projects, use `BaseAgent`, `Environment`, `MultiAgentSimulation` instead — see `references/multi-agent-patterns.md` 4. **Test skeleton:** Unit test stubs, convergence test template, smoke test 5. **Experiment infrastructure:** Config dataclass template, runner script template 6. **Gitignore:** Add `results/`, `*.pkl`, `wandb/`, `__pycache__/` Read `references/algorithm-templates.md` for skeleton code. Read `references/package-scaffold.md` for directory layout. **Gate:** Verify package installs with `uv pip install -e ".[dev]"` before proceeding. ### Phase 3: Experiment Design **Prerequisite:** Working package (Phase 2 or pre-existing). 1. **Config schema:** Python dataclass with validation, defaults, serialization to/from YAML 2. **Config hashing:** SHA-256 fingerprint for reproducibility tracking (see `references/experiment-patterns.md`) 3. **Sweep definitions:** Grid sweep (all combinations), random sweep (budget-limited), manual configs 4. **Seed management:** `np.random.default_rng(seed)` everywhere. Seeds passed through config, never global state. 5. **Runner script:** Config → initialize → loop → collect → save. Reads config, runs `n_seeds` repetitions, saves per-seed results. 6. **Baseline implementations:** Same interface as main algorithm, registered in config 7. **Result aggregation:** Per-seed CSV → aggregated stats (mean ± std). Canonical column naming. 8. **Parallelization:** `concurrent.futures.ProcessPoolExecutor` for independent seeds/configs. **For large sweeps (10+ configs × 10+ seeds, GPU-bound, or >30-min runs):** move to [HPC cluster] HPC — see `docs/guides/hpc.md` in Task Management and copy `templates/slurm/{array,gpu}.sbatch` into `hpc/` with `sync-up.sh` / `sync-down.sh`. Recent reference implementations: `Projects/NLP/{example-project-a,example-project-b}/hpc/`. 9. **Checkpointing:** Save intermediate results to allow resume on crash 10. **Dual output:** Dated archive + "latest" symlink for quick access (see `references/experiment-patterns.md`) **For multi-agent simulations, also include:** 11. **Agent composition:** Heterogeneous population specs with distribution presets (uniform, weighted, custom) 12. **Feature toggles:** Opt-in capabilities per agent type (memory, messaging) — enables ablation studies 13. **Messaging service:** Round-based broadcast channel for inter-agent communication 14. **Multi-level metrics:** Agent-level (individual performance) + system-level (aggregate outcomes) + emergent-level (herding, convergence) See `references/multi-agent-patterns.md` for all multi-agent patterns. Read `references/experiment-patterns.md` for config, sweep, and runner patterns. **Gate:** Run a smoke test — single config, single seed, verify output files are created. ### Phase 3E: Explore Loop (Adaptive Experiments) **Use instead of Phase 3 when:** The design space is unknown, the user wants to adaptively search rather than run a pre-specified sweep, or the user says "explore", "try things", "see what works". Read `references/explore-loop.md` for the full protocol. Summary: 1. **Branch:** Create `experiments/` branch (e.g. `experiments/mar14-collusion-params`) 2. **Baseline:** Run the code as-is, record baseline metric in `results.tsv` 3. **Loop** (autonomous until interrupted or budget exhausted): - Propose a modification based on results so far (informed by prior keep/discard outcomes) - Apply the change to the target file(s) - `git commit` the change - Run with time budget: `timeout s uv run python