--- name: ml-mlip-nvalchemi description: Optional experimental GPU-accelerated batched inference for MACE, MatGL (TensorNet/M3GNet/CHGNet), and FairChem MLIPs using NValchemi, enabling parallel static, relax, and MD workflows across multiple structures simultaneously. metadata: category: [machine-learning] venv: [fairchem, mlip] --- # ml-mlip-nvalchemi ## Goal Run optional GPU-parallel static predictions, relaxation and MD across multiple structures with NValchemi. **This backend is experimental and disabled by default in AtomisticSkills 2.0.0.** List and directory inputs normally run each structure through the model's native calculator and ASE/MatCalc. Installing `nvalchemi-toolkit` or selecting a GPU does not enable the batch engine. > [!WARNING] > The locked toolkit 0.2.0 has confirmed batch-dynamics correctness defects. > Same-size refills and variable-cell changes can omit periodic neighbors; > AtomisticSkills guards this cache defect for 0.2.x, with measurable overhead. > A separate defect can leave live energies inconsistent with frozen geometry > after convergence; preserving the first converged snapshot reduces exposure > but does not repair the upstream live batch. MatGL inflight remains disabled. > Native NPT currently falls back to ASE because initial stress is missing. > Static predictions do not reuse the defective dynamics cache, but this does > not establish correctness for every model or structure. Historical speedups > are workload-specific and do not certify the release's dynamics paths. **Explicit selection:** pass `use_nvalchemi=True` on each batch call to `static_calculation`, `relax_structure` or `run_md`. MCP tools expose the same flag on `predict_structure`, `relax_structure` and `run_md`. The default is `False`; the choice does not persist into later calls. Runtime logs identify experimental use, and batch results report the actual `backend`. When offering this option, explain these limitations. Do not enable it merely because batching could be faster. Validate a small representative case against the default path before a larger run, and check `backend` for any fallback. Single-structure calls continue using their normal calculator. ## Background NValchemi provides batched dynamics integrators (FIRE, NVT Nose-Hoover, NPT, etc.) and a `BaseModelMixin` interface. AtomisticSkills wraps each MLIP in a `BaseModelMixin`-compatible class: | MLIP | NValchemi wrapper | Location | |------|------------------|----------| | MACE | `nvalchemi.models.mace.MACEWrapper` | upstream (nvalchemi-toolkit) | | MatGL TensorNet | `matgl.ext._alchmtk.TensorNetWrapper` | matgl package | | MatGL M3GNet | `M3GNetWrapper` | `src/utils/mlips/nvalchemi/matgl_wrappers.py` | | MatGL CHGNet | `CHGNetWrapper` | `src/utils/mlips/nvalchemi/matgl_wrappers.py` | | MatGL QET | `QETWrapper` | matgl package | | FairChem UMA | `FairChemWrapper` | `src/utils/mlips/nvalchemi/fairchem_nv.py` | With `use_nvalchemi=True`, `src/utils/mlips/base.py` dispatches as follows: - `static_calculation(list, use_nvalchemi=True)` → `_batch_static_nvalchemi()` → single batched forward - `relax_structure(list, use_nvalchemi=True)` → `_batch_relax_nvalchemi()` → batched FIRE - `run_md(list, use_nvalchemi=True)` → `_batch_md_nvalchemi()` → batched NVT/NVE/NPT integrator ### Inflight batching (relaxation) After explicit opt-in, relaxation can choose fixed-batch or inflight execution: ``` _batch_relax() ├─ use_nvalchemi=True AND nvalchemi available AND model loads? │ YES → _batch_relax_nvalchemi() │ └─ sum(atoms) > max_batch_atoms AND model._nvalchemi_supports_inflight? │ YES → _batch_relax_nvalchemi_inflight() ← rolling GPU window │ NO → fixed-batch NValchemi ← all structures at once │ NO → _batch_relax_sequential() ← plain ASE FIRE, one by one ``` > **Note**: All MatGL wrappers (`TensorNetWrapper`, `M3GNetWrapper`, `CHGNetWrapper`) set `_nvalchemi_supports_inflight=False` and use fixed-batch NValchemi regardless of structure count, because after graduation, energies are wrong (TensorNet Cu −83.70 vs −86.57 eV fixed-batch; CHGNet 0.26 eV; M3GNet 28 meV), MACE inflight is available after opt-in with the cache guard. **Inflight batching** keeps only `max_batch_atoms` atoms on the GPU at once. As each structure converges or exhausts its step budget it is evicted and a new one is loaded. This is necessary when the full set of structures would exceed GPU memory. **`max_batch_atoms` — how it is set:** | How | Behaviour | |-----|-----------| | `max_batch_atoms=None` (default) | Auto-estimated: `free_VRAM × 0.5 / bytes_per_atom` using per-architecture calibration (FairChem 0.15 B/param/atom, MACE 0.5, M3GNet 4.0, fallback 5 MB/atom) | | Explicit integer (e.g. `500`) | Forces inflight for almost any real dataset; recommended on shared GPUs | | Very large integer | Forces fixed-batch (all structures in one GPU call) | ### How to tell which backend ran Every batch result dict carries a `"backend"` key: | `"backend"` value | Meaning | |---|---| | `"nvalchemi_inflight"` | Inflight rolling-window (large datasets / shared GPU) | | `"nvalchemi"` | Fixed-batch NValchemi (entire set fits in one GPU pass) | | `"sequential"` | Plain ASE FIRE, one structure at a time | ```python result = wrapper.relax_structure(structures, fmax=0.05, steps=500, use_nvalchemi=True) print(result["backend"]) # "nvalchemi_inflight" / "nvalchemi" / "sequential" ``` Logger messages (written to stderr / MCP server log) also signal transitions: - `"Total atoms (N) exceeds batch limit (M); switching to inflight batching."` - `"NValchemi inflight relax: N structures, live batch ≤M atoms, ≤S steps/structure."` ### Variable-cell relaxation validation Variable-cell relaxation uses upstream `FIRE2VariableCell` with `dt=0.05`, `tmax=0.5`, `delaystep=5`, and `maxstep=0.2`. Its cell force is already normalized by system size; AtomisticSkills no longer applies an additional atom-count scaling. Standard-form cell preparation is retained. All three relaxation modes report `success` only on convergence, `not_converged` at the step limit, and `failed` for execution errors. Results include `converged` and per-structure `steps`, with separate aggregate counts. For variable-cell runs, convergence includes the per-atom virial row norm as well as atomic forces. This is the small-strain counterpart of ASE's `FrechetCellFilter` criterion, not exact equality at finite strain. **Neighbor-cache handling in 2.0.0:** NValchemi 0.2.x can omit neighbors after same-size refills or gradual cell changes. AtomisticSkills uses a version-gated guard for every neighbor hook in fixed relaxation, inflight relaxation and batch MD. Changes to cell, PBC or atom partition trigger a complete allocation refresh; ordinary fixed-cell steps retain their buffers. Installed packages are unchanged. MACE inflight stays available after opt-in, and FairChem builds its own graph without this hook. See the [exposure and timing report](../../docs/verification/nvalchemi-neighbor-cache.md). The separate upstream inactive-output defect can still return live energies inconsistent with frozen coordinates after convergence. MatGL inflight remains disabled pending validation of that repair. Trajectory extraction preserves the first converged snapshot; turning extraction off exposes the live-batch output limitation. A corrected upstream release is still needed to retire the compatibility guard. The former variable-cell speedup tables used force-only convergence and are withdrawn. New speedups remain pending a supported upstream release containing the neighbor-cache and inactive-output repairs. A source-checkout validation run must not be presented as performance of the committed dependency locks. Static-inference and MD tables below measure different operations. ### Molecular Dynamics (MD) Benchmark: Sequential vs. Batched (20 structures, 100 steps) Speedup comparison for a 100-step MD simulation under the `nvt_nose_hoover` ensemble at 300 K on 20 strained Cu FCC structures, each expanded to a fixed 108-atom cubic supercell ($\ge 10\text{ \AA}$ sides). Sequential = NValchemi disabled, structures run one at a time; Batched = all 20 driven through NValchemi integrators in a single GPU batch. Best-of-2 wall time, measured serially (one environment at a time to avoid GPU contention). #### MACE-OMAT-0-small (`mlip`) - **Sequential MD:** 54.48 s - **Batched MD (NValchemi):** 11.12 s (**4.90x speedup**) #### TensorNet-PES-MatPES-PBE-2025.2 (`mlip`) - **Sequential MD:** baseline - **Batched MD (NValchemi):** **1.8× speedup** over sequential (16 structures × 200 steps on GB10; batch NVE energy drift equals sequential, max 0.02 meV/atom) #### FairChem uma-s-1p2 (`fairchem`) - **Sequential MD:** 339.74 s - **Batched MD:** disabled — routed to sequential (see note below; measured ~0.64x, i.e. *slower*, before being disabled) > **When does batched MD help?** Batched MD yields substantial speedups for launch-latency-bound models at small system sizes (e.g. MACE at **4.90x**, TensorNet at **1.8x**). For heavy models like FairChem uma-s-1p2 whose single-system path is already compute-bound, batching provides no speedup and MD is routed to sequential. The wrappers still accept a list of structures (and batch **static/relax** remain available); only FairChem's MD path is gated. Measured on NVIDIA GB10 (aarch64, CUDA 13, Warp 1.14). > **FairChem batched MD disabled (`_nvalchemi_supports_batch_md = False`):** uma-s-1p2's forward scales **superlinearly per atom** — ≈1.57 ms/atom at batch=1 (108 atoms) rising to ≈2.43 ms/atom at batch=20 (2160 atoms), 1.55x worse — so a single large batched step is *slower* than running the structures one at a time through the model's optimized single-system path (batched 0.64x). Two facts pin this down: (1) the cost is intrinsic to the eSCN/MoE forward, not the neighbor list — correcting the wrapper cutoff (12 A → the model's true 6 A) cut `adapt_input` edges from 530 to 78 per atom but left the per-step time unchanged at ~5.25 s; (2) uma-s-1p2 runs with `external_graph_gen=False`, so it **rebuilds its own graph internally and ignores the edges `adapt_input` provides** (energies are identical for any cutoff we pass, including a 0-edge 2 A list). Batched MD is therefore correct (energies match sequential to 0.00 meV/atom) but never a speedup, so `run_md` falls back to sequential. > **TensorNet batched MD enabled with stream fix:** TensorNet batch MD is re-enabled. The `NeighborListHook` "race" it was previously disabled for was a CUDA stream mismatch, now fixed by `warp_on_torch_stream` in `src/utils/mlips/nvalchemi/nvalchemi_utils.py`. Batch NVE energy drift equals sequential (max 0.02 meV/atom), and batch MD is 1.8× faster than sequential for 16 structures × 200 steps on GB10. ## Instructions ### Step 1 — Verify NValchemi is Available ```python # (or venv/fairchem for FairChem models) from src.utils.mlips.nvalchemi.nvalchemi_utils import NVALCHEMI_AVAILABLE print(NVALCHEMI_AVAILABLE) # must be True from src.utils.mlips.mace.mace_wrapper import MACEWrapper wrapper = MACEWrapper(model_name="MACE-OMAT-0-small", device="cuda") wrapper.load() nv = wrapper._get_nvalchemi_model() print(nv) # should be non-None MACEWrapper(nvalchemi) ``` ### Step 2 — Batch Static Calculation Pass a list of ASE Atoms objects to `static_calculation`. The result dict includes a `"backend": "nvalchemi"` key when the batch path was used: ```python from ase.build import bulk import numpy as np structures = [bulk("Cu", "fcc", a=3.6 * s) for s in np.linspace(0.96, 1.04, 10)] result = wrapper.static_calculation(structures, use_nvalchemi=True) # result["backend"] == "nvalchemi" # result["total_structures"] == 10 # result["results"][i] == {"energy": ..., "forces": ..., "stress": ...} ``` Identical API for MatGL and FairChem wrappers: ```python from src.utils.mlips.matgl.matgl_wrapper import MatGLWrapper wrapper = MatGLWrapper(model_name="TensorNet-PES-MatPES-PBE-2025.2", device="cuda") wrapper.load() result = wrapper.static_calculation(structures, use_nvalchemi=True) ``` ```python from src.utils.mlips.fairchem.fairchem_wrapper import FAIRCHEMWrapper wrapper = FAIRCHEMWrapper(model_name="uma-s-1p2", device="cuda") wrapper.load() result = wrapper.static_calculation(structures, use_nvalchemi=True) ``` ### Step 3 — Batch Geometry Relaxation ```python result = wrapper.relax_structure( structure_data=structures, # list of ASE Atoms use_nvalchemi=True, # explicit experimental-backend selection fmax=0.05, # eV/Å convergence steps=500, output_dir="/path/to/output", # relax_cell=True, # optional: variable-cell relaxation # max_batch_atoms=500, # optional: set explicitly on shared GPUs to force # # inflight mode and avoid OOM; None = auto from VRAM ) print(result["backend"]) # "nvalchemi_inflight", "nvalchemi", or "sequential" ``` Variable-cell batch relaxation (`relax_cell=True`) converges only when the per-atom virial row norm is also below `fmax` (the small-strain counterpart of ASE's `FrechetCellFilter`). Per-structure `relax.log` files (ASE FIRE format) are written incrementally to `{output_dir}/{structure_name}/relax.log` during inflight runs, so partial results survive an OOM abort. ### Step 4 — Batch Molecular Dynamics ```python result = wrapper.run_md( structure_data=structures, use_nvalchemi=True, temperature=1000, steps=1000, timestep=2.0, # fs ensemble="nvt_nose_hoover", output_dir="/path/to/output" ) ``` Validated fixed-cell batch families: `nve`, `nvt_nose_hoover`, `nvt_langevin`. NPT aliases select a NValchemi integrator but currently fall back to ASE because the integration does not publish initial stress; native NPT is not validated. Unsupported (Berendsen, Andersen, inhomogeneous NPT) fall back to sequential automatically. ### Step 5 — Use the Default Backend Omit the flag, or set it to `False`; no module patching is needed: ```python result = wrapper.static_calculation(structures) assert result["backend"] == "sequential" result = wrapper.relax_structure(structures, use_nvalchemi=False) ``` MCP/CLI example of an explicit experimental request (two structure files): ```bash ${CLAUDE_SKILL_DIR}/../../venv/run mlip python -m src.mcp_server.cli mace \ load_model model_name=MACE-OMAT-0-small device=cuda \ predict_structure 'structure_data=["first.cif","second.cif"]' use_nvalchemi=true ``` ### Step 6 — Run the Benchmark Script To re-run the full accuracy and speed benchmark for any environment: ```bash ${CLAUDE_SKILL_DIR}/../../venv/run mlip python ${CLAUDE_SKILL_DIR}/scripts/run_nvalchemi_benchmark.py \ --env mace \ --n-repeat 3 \ --output results_mace.json ${CLAUDE_SKILL_DIR}/../../venv/run mlip python ${CLAUDE_SKILL_DIR}/scripts/run_nvalchemi_benchmark.py \ --env matgl \ --n-repeat 3 \ --output results_matgl.json ${CLAUDE_SKILL_DIR}/../../venv/run fairchem python ${CLAUDE_SKILL_DIR}/scripts/run_nvalchemi_benchmark.py \ --env fairchem \ --n-repeat 3 \ --output results_fairchem.json ``` The script tests N=2, 5, 10, 20 structures and prints a speedup/accuracy table. ## Benchmark Results See [resources/benchmark_results.md](resources/benchmark_results.md) for the full results table. ### Speedup Summary (GPU, NVIDIA GB10 Blackwell cc12.1, best-of-3, N=20 structures) | Model | Speedup (N=5) | Speedup (N=20) | ΔE max (eV) | |-------|:---:|:---:|:---:| | MACE-OMAT-0-small | 21.7× | **68×** | 9.5e-07 | | MACE-OMAT-0-medium | 22.9× | **72×** | 9.5e-07 | | MACE-MH-1/omat_pbe | 14.3× | **34×** | 1.6e-07 | | MACE-MH-1/matpes_r2scan | 14.4× | **34×** | 1.2e-07 | | MACE-MP-medium-0b3 | 22.9× | **76×** | 1.4e-06 | | MACE-MATPES-PBE-0 | 23.6× | **77×** | 7.2e-07 | | MACE-MATPES-R2SCAN-0 | 24.2× | **76×** | 1.9e-06 | | TensorNet-PES-MatPES-PBE-2025.2 | 3.6× | **11×** | 1.4e-07 | | TensorNet-PES-MatPES-r2SCAN-2025.2 | 3.8× | **12×** | 8.0e-07 | | M3GNet-PES-MatPES-PBE-2025.2 | 3.7× | **11×** | 1.3e-03¹ | | M3GNet-PES-MatPES-r2SCAN-2025.2 | 3.9× | **11×** | 7.2e-04¹ | | CHGNet-PES-MatPES-PBE-1M-2026.9 | 4.3× | **12×** | 2.4e-07 | | CHGNet-PES-MatPES-r2SCAN-1M-2026.9 | 4.2× | **13×** | 9.5e-07 | | QET-PES-MatPES-PBE-2025.2 | 4.2× | **13×** | 7.2e-07 | | QET-PES-MatPES-r2SCAN-2025.2 | 3.9× | **14×** | 9.5e-07 | | SO3Net-PES-ANI-1x-Subset | — | — | not supported | | FairChem uma-s-1p2 (omat) | 3.0× | 2.9× | 1.6e-07 | | FairChem uma-m-1p1 (omat) | 2.9× | 3.5× | 1.7e-07 | | FairChem uma-s-1p1 (omat) | 3.4× | 5.5× | 2.5e-07 | ¹ M3GNet energy errors (~0.7–1.8×10⁻³ eV) from different neighbor-list graph connectivity (NValchemi GPU warp kernel vs. CPU `radius_graph_pbc`). Forces are exact (ΔF = 0). Within 5×10⁻³ eV tolerance for PES screening. ## Constraints - **Explicit opt-in required**: Set `use_nvalchemi=True` for each batch request. - **NValchemi required**: `nvalchemi-toolkit` must be installed. Check `NVALCHEMI_AVAILABLE` flag. Falls back to sequential if unavailable. - **Environment isolation**: Must use the correct uv environment per MLIP: - `mlip` — MACE models and MatGL (TensorNet, M3GNet, CHGNet) - `fairchem` — FairChem UMA - **Stress format**: NValchemi returns 3×3 Cauchy stress tensor (eV/ų); sequential path returns ASE Voigt-6. Both formats are accepted downstream — `_extract_static()` in tests handles the conversion. - **FairChem dataset field**: UMA model requires `dataset` (e.g., `"omat"`) passed to `FCAtomicData`. This is handled automatically by `FairChemWrapper`; defaults to `"omat"` when `task_name=None`. - **CHGNet batch speedup**: CHGNet directed line graph construction parallelizes well on GPU (12–13× at N=20). CPU performance is marginal (<3×); always use `device="cuda"` for batch workloads. - **SO3Net not supported**: `SO3Net-PES-ANI-1x-Subset` falls back to sequential automatically (`_get_nvalchemi_model()` returns `None`). - **ANI-1x models with transition metals**: TensorNet-PES-ANI-1x and M3GNet-PES-ANI-1x training sets cover only H/C/N/O. Using them with Cu or other transition metals causes a CUDA index OOB error that corrupts the CUDA context for the session. Run ANI-1x models in a separate process from other models. - **MatGL models (TensorNet, CHGNet, M3GNet) inflight batching not supported**: Inflight batching stays off for the MatGL wrappers (TensorNet, M3GNet, CHGNet), for a measured reason: after graduation, energies are wrong (TensorNet Cu −83.70 vs −86.57 eV fixed-batch; CHGNet 0.26 eV; M3GNet 28 meV), while MACE inflight agrees to meV. All MatGL wrappers set `_nvalchemi_supports_inflight=False`; when the total atom count exceeds the batch budget, they fall through to fixed-batch NValchemi (all structures in one GPU pass) rather than inflight. For large MatGL sets, split inputs into smaller calls or retain default sequential execution. `max_batch_atoms` does not cap the fixed-batch allocation when inflight is disabled. - **Unsupported ensembles for batch MD**: `nvt_berendsen`, `nvt_andersen`, `nvt_bussi`, `npt_berendsen`, and `npt_inhomogeneous` have no NValchemi equivalent and always run sequentially. ## References - NValchemi toolkit: NVIDIA internal package (nvalchemi-toolkit, PyPI: `https://pypi.nvidia.com`) - MACE: Batatia et al., "MACE: Higher Order Equivariant Message Passing Neural Networks for Fast and Accurate Force Fields", *NeurIPS 2022*. [arXiv:2206.07697](https://arxiv.org/abs/2206.07697) - MatGL / TensorNet: Chen & Ong, "A Universal Graph Deep Learning Interatomic Potential for the Periodic Table", *Nature Computational Science 2023*. [DOI:10.1038/s43588-022-00349-3](https://doi.org/10.1038/s43588-022-00349-3) - M3GNet: Chen & Ong, "A universal graph deep learning interatomic potential for the periodic table", *Nature Computational Science 2022*. - CHGNet: Deng et al., "CHGNet as a pretrained universal neural network potential for charge-informed atomistic modelling", *Nature Machine Intelligence 2023*. [DOI:10.1038/s42256-023-00716-3](https://doi.org/10.1038/s42256-023-00716-3) - FairChem UMA: Meta FAIR, "Scaling Universal Molecular Atomistic Machine Learning for Open Catalyst 2024". [arXiv:2411.12234](https://arxiv.org/abs/2411.12234) --- **Author:** Bowen Deng **Contact:** [github.com/bowen-bd](https://github.com/bowen-bd)