--- name: jcm-local-ci description: Run the jax-gcm CI gates locally on Derecho when GitHub Actions minutes are exhausted or a pre-push check is wanted — lint, fast tests (90% coverage), slow tests (80% PR coverage) and a local Claude code review. Use before pushing or merging any jcm branch without Actions. --- # Local CI for jax-gcm Reproduces the CI gates from `.github/workflows/run_test.yaml` + `run_linter.yaml` (lint, fast, slow and extras), plus the Claude review, without GitHub Actions. ## One command ```bash scripts/local_ci.sh /path/to/worktree # lint here, both gates via qsub scripts/local_ci.sh --local-fast /path/to/worktree # ...and a fast gate on this node ``` Lint runs on the current node; both test gates run in a single `develop`-queue PBS job (`select=1:ncpus=16:mem=200GB`), fast then slow, sequentially. Watch the job log for `FAST_EXIT=0`, `SLOW_EXIT=0` and the closing `GATES PASSED`: the job exits non-zero if either gate failed, so a job that ends green means both passed. Every submission gets its own log, named for the tree it tests: `/jcm_ci...log`, where `` is HEAD's short sha, with `+dirty.` appended when the worktree differs from HEAD — tracked edits or untracked, non-ignored files, which pytest would collect too; the hash is of that difference (the tracked diff and every untracked file's contents), so two different dirty trees on one commit get different tags. The PBS job is named `jcm_ci__` to match (punctuation replaced by `_`, to keep the name to characters every PBS accepts). The script prints the exact log path and the watch command when it submits: ```bash tree: 1a2b3c4d.20260930T051200Z log: /path/to/worktree/jcm_ci.1a2b3c4d.20260930T051200Z.log watch: grep -E 'FAST_EXIT|SLOW_EXIT|GATES' /path/to/worktree/jcm_ci.1a2b3c4d.20260930T051200Z.log ``` The job tests the worktree as it is when the job *starts*, which an edit or a commit made while it queued can change, so it recomputes the tag then: the log opens with `submitted= tested=` (and a `WARNING` line when they differ), and every result line carries both, `GATES PASSED tree=. tested=`. Use the printed path, not a glob over `jcm_ci.*.log`: an earlier run's log in the same worktree is a verdict on an earlier tree, and a green line in it says nothing about the one just submitted. Check that the `tested=` tag on the line you read is the sha you pushed. The gates go to a compute node because a login node caps you at 10 GiB (see `docs/source/design/test_suite_memory.md`) — well under what an `-n 12` fast suite needs, so a local run there reports OOM-killed workers as unrelated test failures. `--local-fast` opts into an `-n 2` fast run on the current node anyway, for a quick read before the job lands; it runs *before* the submission, never alongside it, because both would fight over the same worktree's `.coverage.*`. ## Which dinosaur the gate tests against The gate exists to reproduce CI, so it must run the dinosaur that `pip install -e .` resolves — `requirements.txt` pins `dinosaur>=1.5.0`, which carries the semi-Lagrangian transport jcm's backend requires (neuralgcm/dinosaur#135). The script therefore **auto-detects nothing**: a stale fork checkout sitting in `$HOME` would silently displace the pinned package and the gate would measure a dependency CI never sees, which is the one thing it is for. A fork is used only when you pass `JCM_DINOSAUR` explicitly, and the run then says so and warns that it is *not* at CI parity. With no override and an installed dinosaur that lacks the SL class, the script fails before lint and names both remedies — `pip install -e .`, or `JCM_DINOSAUR=`. (Failing is the point: without SL every model-construction test raises, and ~100 unrelated failures bury the ones that matter.) `~/.venvs/jaxgcm` satisfies the pin (dinosaur 1.5.0), so gate jobs from it need no override — just run the script. `JCM_DINOSAUR` remains the escape hatch for a pre-release checkout, and a run using it says so and states that it is not at CI dependency parity. ## The gates, individually ```bash # 1. Lint (run_linter.yaml) ruff check . # 2. Push gate — fast tests, 90% coverage JAX_PLATFORMS=cpu pytest -n 12 -m "not slow" --cov=jcm --cov-fail-under=90 coverage report --fail-under=90 # 3. PR gate — slow tests only, 80% coverage vs .coveragerc-pr JAX_PLATFORMS=cpu pytest -n 4 -m "slow" --cov=jcm \ --cov-config=.coveragerc-pr --cov-fail-under=80 coverage report --rcfile=.coveragerc-pr --fail-under=80 ``` ```bash # 4. extras-tests — every test an optional extra gates, fast and slow, in a # SEPARATE venv (installing the extras into the coverage venv breaks # dependency parity for gates 2 and 3): pip install -e ".[$(python tools/ci/optional_extras.py pip-extras)]" python tools/ci/optional_extras.py check JCM_REQUIRE_EXTRAS=1 JAX_PLATFORMS=cpu pytest -v -rs -m requires_extra ``` Gate 4 mirrors the `extras-tests` job. Under `JCM_REQUIRE_EXTRAS=1` the session refuses to start without every extra and a selected test that skips is a failure, so a green run means every gated test ran. It is serial on purpose, as in CI: the pySES delegating-config chunk peaks at ~12.7 GB. About 13 min on the dev workstation. A test gates on an extra only through `@pytest.mark.requires_extra(...)`; any other gate fails gates 2/3 (see `tools/ci/optional_extras.py`). `local_ci.sh` does not run it: it needs its own venv, so run the three commands yourself. Gate 3 runs the whole slow suite in one command. CI runs the same tests as two parallel path shards (`slow-tests-radiation` / `slow-tests-rest`, defined with their partition check in `tools/ci/slow_shards.py`) and enforces the 80% floor on their combined coverage in `slow-coverage`; the local gate is equivalent. The trailing `coverage report` is not redundant — it mirrors the two enforcement steps CI gained in #786, and it is the one that exits non-zero on plugin behaviour nobody controls. See the "Coverage differs by suite" note below. Gates 2 and 3 are real compute — run them on a `develop`-queue node, not a login node; `local_ci.sh` does. The `-n 4` on gate 3 is **mandatory on Derecho**, not an optimization: a single serial process accumulates thousands of mmap'd XLA JIT code sections and dies mid-suite with `LLVM ERROR: Unable to allocate section memory!` (SIGSEGV/SIGABRT — three attempts at 120–220 GB all failed identically; RAM is not the issue, per-process map count is). Splitting across workers resets the budget. GitHub's runners tolerate the serial run; Derecho's do not. (`docs/source/design/test_suite_memory.md` covers the related growth in retained XLA executables that the root `conftest.py` bounds.) ## Local Claude review — run it BEFORE pushing, on the right model The adversarial self-review `jcm-dev-workflow` step 3 requires before every push that changes code: Codex credits are finite, and a finding Codex makes that this review would have made is a credit burnt and a round lost. Fix or explicitly refute every finding before `git push`. **Which reviewer depends on the session's model.** `/code-review high` forks the invoking session — same model, full conversation context inherited — and its finders and verifiers fan out the same way, so it is billed to *that* session at full context. It exhausted a Fable session limit twice on 2026-09-10 (the failure notices name the model: `model sent to the API: claude-fable-5-1`). - **Opus or Sonnet session:** `/code-review high` on the branch, as before. - **Fable session:** do not invoke the skill. Spawn a general-purpose agent with `model: opus`, give it the same scope (reference formulation, JAX hygiene, tests, docs, comment style, diff vs upstream), and have it post one review via `gh pr review --comment --body-file `. Forward its findings to the authoring agent for fix-and-reply. This is how #776's blocker and #783's second-round findings were caught. - If the maintainer explicitly asks for `/code-review` from a Fable session, say first that it will fork on Fable at full context, and confirm. For the deep cloud variant use `/code-review ultra` (user-triggered, separately billed). ## Codex review comments: always reply inline Every Codex (or other bot) inline comment on a PR gets an **explicit threaded reply** stating the resolution — "Confirmed and fixed in : " or "Refuted: " — via ```bash gh api -X POST repos///pulls//comments//replies \ -f body="..." ``` (comment ids from `gh api repos///pulls//comments`). Fixing the code silently is not enough: the reviewer tracks resolution through the comment threads, and an unanswered thread reads as an unaddressed finding. Verify a claim against the actual code/data before replying — Codex has been right (forcing unit conventions) and wrong (FZJ ozone units) on the same PR. ## Facts that bite - **Login-node load produces phantom failures.** A `-n 12` fast-suite run on a busy login node has produced 50+ failures across unrelated subsystems that all pass serially (twice now: 55 in July, 53 in August). That is why the fast gate is a PBS job too. Before believing a red `--local-fast` run, rerun a sample of the failures serially. - **Never run two coverage suites concurrently in one worktree.** pytest-cov erases `.coverage.*` at startup and combines at exit, so a fast-gate run (or a stray `rm .coverage*`) deletes an overlapping slow job's in-flight worker data — all its tests pass but whole modules lose credit (76%, then 0.00%, on a tree whose true number was 85%). Sequence the gates, or give the slow PBS job its own worktree. - **Capture pytest's exit via `PIPESTATUS[0]`**, never `$?` after a `| tail` — and never pipe the slow suite through `tail` at all (a segfault's context ends up truncated). - **Lint the whole repo (`ruff check .`), never a subdirectory** — CI lints everything, including tools/ and stray root files. And never commit with `git add -A`: it sweeps untracked scratch files into the commit, which is exactly how a lint-clean `jcm/` still turned CI red once. Stage files by name. - **Coverage differs by suite.** The fast gate uses `.coveragerc` (builders omitted); the slow gate uses `.coveragerc-pr` (also omits fast-tested utility modules). New fast-tested modules that slow tests never touch belong in `.coveragerc-pr`'s omit list — that is the repo's documented mechanism, see the header comment there. - **A floor is only enforced at the reported precision.** `fail_under` is checked as `round(total, precision) < fail_under`, so at coverage's default precision of 0 the 80 floor was a 79.5 floor and the slow gate printed `FAIL ... 79.68%` while exiting 0 (#786). Both rcfiles now carry `[report] precision = 2`; run the gates with the repo's rcfiles (never a bare `--cov-fail-under` against a config that lacks it), and keep the standalone `coverage report` line so a plugin change cannot disarm the gate unnoticed. - `JAX_PLATFORMS=cpu` is mandatory on GPU nodes (xdist workers otherwise fight over the GPU). - The gate job asks for 16 cpus so `-n 12` (fast) and `-n 4` (slow) both fit, and 200 GB because the suite is memory-bound, not CPU-bound. - `local_ci.sh` exports `JAX_COMPILATION_CACHE_DIR` (default `$SCRATCH/jcm-jax-cache`) so xdist workers and successive gate runs share XLA compiles of identical jitted modules instead of each recompiling. Only the XLA-compile share is saved — tracing reruns — and only for bit-identical whole model steps. Safe: a miss just recompiles. - The repo pins `ruff` in CI, in two places that must agree — `run_linter.yaml` (lints everything, every push) and the `lint` gate job in `run_test.yaml` (the fast/slow suites hang off it). Run the pinned version (`pip show ruff` vs either file) before trusting a clean pass. - GPU-gated slow tests skip on CPU exactly as they do in CI — a local CPU pass is equivalent evidence.