--- name: jcm-benchmark description: Measure jcm throughput reproducibly — short (1 month) or long (12 month) runs on a validated stable config, with GPU memory/utilisation logging and an explicit convergence criterion. Use when comparing hardware, library versions (jax-rrtmgp, dinosaur), physics presets, or checking for a performance regression. --- # Benchmarking jcm Builds on `jcm-run` (config groups, Hydra traps) and the machine skill for your site — `devbox-jcm-runs` or `derecho-jcm-runs`. This skill is about getting a throughput number that is *actually true*, which is harder than it looks. ## Run it ```bash PY=/home/dwatsonparris/micromamba/envs/jcm/bin/python # Short: 1 month (30 d), ~6 chunks. The default for any A/B. $PY tools/benchmark.py --preset t63-echam-rrtmgp --months 1 --gpu 1 \ --label baseline # Long: 12 months. Use for provisioning and drift, not for A/B. $PY tools/benchmark.py --preset t63-echam-jam --months 12 --gpu 3 \ --chunk-days 30 --label jam-year # A/B a library version without touching the shared editable install $PY tools/benchmark.py --preset t63-echam-rrtmgp --months 1 --gpu 1 \ --label rrtmgp-perf \ --pythonpath /data/dwatsonparris/jax-rrtmgp-worktrees/perf ``` Presets are in `tools/benchmark.py:PRESETS`. Each carries the **validated stable override set** for its grid, not just `physics=`/`grid=` — see "Configuration" below. Results land in `/scr/dwatsonparris/benchmarks/