--- name: jcm-run description: Launch a jcm model run through the built-in Hydra configs — config groups, the validated stable T63L47 overrides, Hydra override traps, and watching for startup failures. Site-agnostic; pair with devbox-jcm-runs (shared workstation) or derecho-jcm-runs (NCAR PBS) for machine specifics. --- # Running jcm **Layering.** This skill is the site-agnostic model layer. For where to run, see the machine skill: `devbox-jcm-runs` (shared UCSD workstation, no scheduler, you pick the GPU) or `derecho-jcm-runs` (NCAR Derecho, PBS allocates GPUs). For throughput measurement see `jcm-benchmark`. Every runnable configuration goes through `python -m jcm.main` with Hydra groups and overrides. **Never write a bespoke driver script** — see the "No bespoke run scripts" rule in `CLAUDE.md`. If a configuration is worth repeating, it becomes a config file under `jcm/config//`. ## Config groups | group | options | |---|---| | `physics` | `speedy`, `held_suarez`, `echam`, `echam-rrtmgp`, `echam-rrtmgp-2m`, `echam-rrtmgp-2m-cosp`, `echam-strong-conv`, `echam-jam`, `echam-jam-aerocom`, `echam-jam-aerocom-optics`, `echam-jam-aci` | | `grid` | `speedy_t31_l8`, `held_suarez_t31_l8`, `echam_t42_l8_sigma`, `echam_t63_l47_hybrid`, `echam_t85_l47_hybrid`, `echam_t63_l95_hybrid`, `echam_t106_l95_hybrid`, `echam_t119_l95_hybrid` | | `init` | `isothermal`, `balanced_isothermal`, `jw` | | `terrain` | `aquaplanet`, `from_file` | | `forcing` | `default`, `from_file` | | `run` | `default`, `smoke`, `longrun`, `pyses_year` | | `dycore` | `dinosaur`, `pyses_ne30l47` | | `diffusion` | `default`, `strong` | Discover current options rather than trusting this table if it looks stale: `ls jcm/config//`. ## The validated T63L47 ECHAM launch This is the known-stable production baseline. **Use it as the starting point for any T63L47 run** — the pieces below are not optional decoration, they are what keeps the run from going NaN (see "Stability" below). ```bash PY=/home/dwatsonparris/micromamba/envs/jcm/bin/python REPO=/data/dwatsonparris/jax-gcm TS=$(date +%y%m%d_%H%M%S) COMMON="physics=echam \ grid=echam_t63_l47_hybrid \ init=jw init.rh=0.0 \ terrain=from_file terrain.file=hf://bundles/t63/terrain.nc \ forcing=from_file forcing.file=hf://bundles/t63/forcing_pd.nc \ run=longrun" PREFIX=myrun_$TS nohup env CUDA_VISIBLE_DEVICES=0 XLA_PYTHON_CLIENT_PREALLOCATE=false \ $PY -m jcm.main $COMMON \ run.output_prefix=$PREFIX \ +run.checkpoint_path=${PREFIX}.ckpt \ > $REPO/run_logs/${PREFIX}.log 2>&1 & echo "$PREFIX PID=$!" ``` Write outputs to `/scr/dwatsonparris/...` for anything large — `/data` is near-full. Logs conventionally go in `run_logs/` at the repo root. ## Pre-flight: where to run GPU selection is machine-specific and lives in the site skills: - **`devbox-jcm-runs`** — shared workstation. You self-allocate, so you must verify a card is genuinely free (`python tools/gpu_util.py`) and avoid stomping on colleagues. Utilisation and memory each *individually* look idle for a parked job; both must be checked. - **`derecho-jcm-runs`** — PBS allocates exclusive GPUs; pre-flight the Hydra composition before spending a queue slot instead. - **`kubernetes-jcm-runs`** — a pod gets exclusive GPUs, so timings are trustworthy and a sweep runs in parallel; in exchange the run must survive eviction and the code has to be pinned by SHA. `JAX_PLATFORMS=cpu` is for **unit tests only**. Anything beyond ~5 simulated days belongs on a GPU. ## Stability: why the overrides matter A T63L47 ECHAM run started from an isothermal cold start with no sponge **will go NaN within a few days**. The stable recipe needs: - `init=jw init.rh=0.0` — Jablonowski-Williamson balanced initial state. `init=isothermal` on a real-orography grid is not a viable start. - `terrain=from_file` + `forcing=from_file` — real orography/land-sea mask and SSTs. - `run=longrun` — one **calendar year** (`run.total_time: 12 months` from `run.start_time`) written as calendar-month means (`_monthly_YYYY-MM.nc`, `run.monthly_means`, #901) from daily means in 5-day chunks; no per-chunk `_dayN.nc` unless `run.save_chunks=true`. `forcing_pd.nc` + the `auto` emission/ozone/oxidant bundles are the present-day (2005–2014) climatological AMIP forcing. - `run=longrun` — this already carries **ECHAM's upper sponge** (`uspnge`: the zonal anomalies of u, v and T at the top level damped on 3 h, the zonal mean untouched; rationale in `run/longrun.yaml`). Do **not** re-specify it on the command line. There is no absolute temperature target any more: `run.sponge.target_T_K` is not a key and an override of it fails. ## Hydra gotchas - **`run.time_step` is in MINUTES**, not seconds. - **`save_interval` must be ≤ `chunk_days`.** Otherwise a chunk contains zero output times and the chunk write dies with a confusing `IndexError: index 0 is out of bounds for axis 0 with size 0` from `predictions.to_xarray()`. This is easy to hit when shortening a run for a quick test and forgetting to shorten `save_interval` with it. - `run/longrun.yaml` has **no** `checkpoint_path` key, so adding one needs Hydra's add syntax: `+run.checkpoint_path=...`. Plain `run.checkpoint_path=...` fails with `Key 'checkpoint_path' is not in struct`. `run/default.yaml` does define it. - **Ozone**: `forcing.ozone_file: auto` is the shipped default and resolves a packaged climatology matching the grid (`jcm/data/bc/t63/ozone.nc` — already on L47 levels, already S→N). Leave it alone. Confirm in the log: `forcing.ozone_file=auto resolved to .../t63/ozone.nc`. On a hybrid grid `auto` now **raises** rather than degrading if it resolves nothing: the analytic profile carries ~7.6× the tropospheric ozone column, a large clear-sky OLR bias and not a valid basis for any radiation comparison. The error distinguishes a missing product from a cold cache and names the remedy. For any other grid `auto` also consults the HF mirror's level-resolved bundles; set one explicitly with `forcing.ozone_file=hf://bundles/_l/ozone_pd.nc` (prefetch on a node with internet), or regenerate a packaged file with `jcm.data.bc.interpolate_ozone` for offline work. `forcing.ozone_file=analytic` takes the analytic profile deliberately; a **sigma** grid, for which no ozone product exists, still warns and falls back. ## Watching a run One `tail -F` with an alternation that covers **both** progress and every failure signature — a filter that only matches success is silent through a crash, which reads identically to "still running": ```bash tail -F -n 0 run_logs/PREFIX.log 2>/dev/null | grep -E --line-buffered \ "Saved .*_day[0-9]+\.nc|Wall: |NaN vars|unhealthy|Traceback|Error|FAILED|Killed|OOM|CUDA_ERROR|HydraException" ``` `NaN vars: N/239` is the health-check line. Parse the count — do **not** grep for the bare string `nan`, which matches unrelated output. `specific_humidity` in the saved netCDF and the health report is genuinely in **g/kg**, and the `units: 'g/kg'` label is correct — `state_bridge` calls `dimensionalize(q, gram/kilogram)` on the way out. Healthy tropical surface values are ~20-30 g/kg. (An earlier version of `check_health` assumed kg/kg and applied a `*1000`, which double-counted and tripped the `q_max > 100` threshold on every chunked run; that conversion was removed. Do not reintroduce the assumption in either direction.) ## Failure modes worth recognising - **Chunk write crashes after "Run completed"** — a diagnostic emitted a shape `data_to_xarray` has no dims for. The write order is `to_xarray → check_health → to_netcdf → save_checkpoint`, so a crash here loses the whole chunk with no checkpoint. Fix by adding the dotted key to `ComposablePhysics._EXCLUDED_OUTPUT_KEYS` or registering a band coord. - **Editable installs**: `jcm`, `jax-rrtmgp` and `mam4-jax` are installed editable, so **the working tree is the running code**. Check `git -C rev-parse --abbrev-ref HEAD` before trusting any result. The site skills cover how to A/B a library version safely (worktree + `PYTHONPATH`, never `git checkout` in a shared clone).