--- name: tao-run-on-slurm description: Remote SLURM GPU cluster execution over SSH with sbatch/srun, Pyxis/Enroot containers, and Lustre-backed results. Use when running TAO training/eval/inference jobs on an on-prem or DGX SLURM cluster. Trigger phrases include "run on SLURM", "submit sbatch", "DGX SLURM cluster", "Pyxis/Enroot container", "Lustre dataset". license: Apache-2.0 compatibility: Requires SSH access to a SLURM login node (passwordless via key auth) and SLURM_USER + SLURM_HOSTNAME env vars. No nvidia-tao-sdk install is required; jobs are driven directly over ssh + sbatch/squeue/sacct/scancel. metadata: author: NVIDIA Corporation version: "0.1.1" allowed-tools: Read Bash tags: - platform - slurm --- # SLURM > **Standalone install?** If this session was not initialized by the TAO skill bank plugin, run the `tao-setup` skill first (host preflight, credentials, cross-skill discovery). Remote GPU compute platform for clusters managed by SLURM. Jobs are submitted from the launch host to a login node over SSH, staged on a shared filesystem, submitted with `sbatch`, and executed with `srun` container support. ## When to use Use SLURM when the user has access to a managed GPU cluster, shared Lustre storage, and scheduler-owned GPU allocation. Do not use SLURM for local files that exist only on the agent machine; data and outputs must be reachable from the cluster. ## Preflight + SSH Confirm `SLURM_USER` and `SLURM_HOSTNAME` are exported and passwordless SSH to a login host works (`ssh -o BatchMode=yes`). The launch host needs `ssh`, not local `sbatch`, `srun`, Enroot, or a Lustre mount. Preflight those scheduler, Pyxis, Enroot, and shared-storage dependencies on the selected remote login/compute frame. Model-specific inspectors may be streamed from the installed skill over SSH stdin; do not stage an ad-hoc source patch or treat the launch host as the SLURM frame. For private `nvcr.io` images, install `~/.config/enroot/.credentials` on the cluster once per (cluster, user): Pyxis/Enroot does not read `NGC_KEY` from the job env, and without persistent credentials, auth-gated pulls fail with "Could not process JSON input" at job startup. Install it via the `printf | ssh` heredoc so the `NGC_KEY` value never lands in shell history, intermediate files, or chat output; never `cat`/`echo` the value. If a preflight check fails, the agent prompts the user to authorize the install/fix via Bash. Pip-installable Python requirements are the exception: install them automatically, then rerun preflight. See `references/slurm-ssh-credentials.md` for the full preflight script, the enroot-credentials heredoc, prerequisite key setup (keypair, `ssh-copy-id`, `known_hosts`, container key mounts, 2FA handling), and the SSH failure remediation prompt. ## Execution — the four verbs `tao-run-on-slurm` is a platform **consumer**: it runs a spec-bundle over `ssh + sbatch/squeue/sacct/scancel`, mutating only the job-record. Storage is **tier A** (Lustre) — the dataset is staged to a shared path *before* submit and read through Pyxis; never fetch S3 inside the allocation (the scheduler-idle timeout kills GPU-idle jobs and bills the wasted time). `$BANK` = `${TAO_SKILL_BANK_PATH}`; `$LOGIN` = a resolved `SLURM_HOSTNAME`. ### submit 1. **Reuse what's already staged — never redo (tier A):** - *Image:* `@@IMAGE@@` is a Lustre `.sqsh` — **reuse an existing one if present** (`ssh $LOGIN ls `); only if missing, convert once with `enroot import` (cached by name — see `references/slurm-container-execution.md`). - *Dataset:* **confirm it is already on Lustre** (`ssh $LOGIN test -e …`) and reference those paths; `tao-data-io` stages *only* a small auxiliary input that is not there yet — never re-stage existing data, and never the training set inside the allocation. Then author the spec at `/specs/spec.yaml` on Lustre with those paths. 2. **Credentials → sidecar (never inline):** if the run needs session creds (e.g. `HF_TOKEN`), write them to a mode-600 sidecar on Lustre and let the template shred it on exit; NGC image pulls use the one-time `~/.config/enroot/.credentials` (see `references/slurm-ssh-credentials.md`), not the job env: ```bash set -a; source /path/to/.env; set +a # omit if already exported printf 'export HF_TOKEN=%s\n' "$HF_TOKEN" | ssh $LOGIN "umask 077; cat > /job_$JOB_ID.env" ``` 3. **Open the record — mints the id, binds `results_dir` on Lustre, before launch:** ```bash JOB_ID=$("$BANK/scripts/tao_job_record.py" open --platform slurm --image "$IMAGE" \ --network-arch "$ARCH" --action "$ACTION" --storage-tier A --results-root "$SLURM_BASE_RESULTS_DIR") ``` 4. **Consume the optional model lifecycle.** If the validated spec-bundle has `execution`, preserve its order and semantics while mapping distributed intent to native SLURM/Pyxis. Stage only its checksum-closed `supporting_files`. The full generic lifecycle and staging contract is in `references/slurm-container-execution.md`. 5. **Render** `templates/slurm/singlenode.sbatch.tmpl` — substitute every `@@@@` (`JOB_NAME=$JOB_ID`, `NUM_GPUS`, `CPUS_PER_TASK`, `TIME`, `LOG_DIR`, `IMAGE`, `CONTAINER_MOUNTS=`, `COMMAND=`, `SBATCH_EXTRA=` account/partition lines, `ENV_FILE=` the sidecar path or empty, `EXTRA_ENV=` any cluster NCCL knobs) → `/sbatch/job_$JOB_ID.sbatch`. **Lint + syntax-check before submit:** `redact_secrets.py lint ` must pass and `bash -n ` must succeed. 6. **Submit + record RUNNING:** ```bash SLURM_ID=$(ssh $LOGIN "sbatch --parsable /sbatch/job_$JOB_ID.sbatch") "$BANK/scripts/tao_job_record.py" mark "$JOB_ID" --state RUNNING --backend-ref "$SLURM_ID" ``` A submit that skipped the gate or the open has no id — so it cannot launch. ### status ```bash # sacct ANNOTATES states ("CANCELLED by 12345") and truncates them to the # default column width, so a cancelled job reads back as "CANCELLED+" and # matches nothing in the table below — reporting UNKNOWN instead of CANCELED. # Widen the column, take the first word, drop the truncation marker. st=$(ssh $LOGIN "sacct -j $SLURM_ID -X -n -o State%30" | awk '{print $1}' | tr -d '+') # (use squeue while the job is still PENDING; sacct lags briefly after submit) ``` | SLURM state | vocab | |---|---| | `PENDING` | `PENDING` | | `RUNNING` / `COMPLETING` | `RUNNING` | | `COMPLETED` | `COMPLETE` (confirm `status.json` in `results_dir`) | | `FAILED` / `TIMEOUT` / `OUT_OF_MEMORY` | `ERROR` (infra-vs-program classify → retry, M6) | | `NODE_FAIL` / `BOOT_FAIL` | `ERROR`, `err_class=ERR_INFRA` (`--requeue` re-queues these) | | `CANCELLED` / `PREEMPTED` / `REVOKED` | `CANCELED` | | (not found) | `UNKNOWN` | Native sub-state rides in the transition `message`. Poll at the chosen interval; long queue waits are normal — do not stop on elapsed time. ### logs ```bash ssh $LOGIN "tail -n ${N:-200} /$JOB_ID-$SLURM_ID/main.out" # SLURM auto-creates the %x-%j subdir ``` ### cancel ```bash ssh $LOGIN "scancel $SLURM_ID" "$BANK/scripts/tao_job_record.py" mark "$JOB_ID" --state CANCELED --source agent ``` Treat an already-terminated SLURM job as a successful cancel. ### Multi-node (nodes > 1) Same four verbs, with three additions at submit: 1. **Render `templates/slurm/multinode.sbatch.tmpl`** instead of the single-node one — it's a strict superset (adds `--nodes` / `--wait-all-nodes` + the rendezvous block). `WORLD_SIZE` is the **node count** (TAO's misnomer); never change it to a global-rank count. 2. **NCCL probe first** — before the real job, run a cheap 2-node all-reduce (`scripts/nccl_allreduce_probe.py` under the container's torchrun) with a ~120s timeout. Before invoking torchrun, preserve the TAO rendezvous values as `TAO_NODE_COUNT=$WORLD_SIZE`, `TAO_GPUS_PER_NODE=$NUM_GPU_PER_NODE`, and `TAO_NODE_RANK=$SLURM_PROCID`; torchrun overwrites its standard `WORLD_SIZE` with the global process count. `NCCL_PROBE_OK` → proceed. **Timed out** (the collective hung) → set the cluster's NCCL knob in `EXTRA_ENV` and re-probe — on CS-OCI-ORD that is `export NCCL_P2P_DISABLE=1` (the intra-node P2P hang), often with `NCCL_SOCKET_IFNAME=eth0` / `NCCL_IB_DISABLE=1`. **Cache the working env per cluster** so later jobs skip the probe. Gate on **`gpus_per_node > 1` too** — the P2P hang triggers on a single node with 2+ GPUs. 3. Tier-A Lustre, sidecar creds, record, and lint are unchanged. ### Cosmos backend guardrails Read [`references/cosmos-slurm-guardrails.md`](references/cosmos-slurm-guardrails.md) before rendering a Cosmos command. It defines image staging, planner materialization, Framework and Cosmos-RL launch contracts, worker/runtime requirements, and exit/status handling. ## Storage Use shared-filesystem URIs, not local or `file://` paths; `tao-core` rejects local/file paths for remote backends. - `lustre:///absolute/path` for user-provided datasets on Lustre. - `slurm://` paths may appear in microservices metadata and are converted to Lustre paths before the container starts. Accept either dataset roots (model skills map them to required files) or direct spec-key paths. After SSH succeeds and before generating scripts, `test -e` each required dataset path from the login host; if it fails, stop and ask for corrected paths or staged data rather than producing scripts that fail in the first training job. See `references/slurm-ssh-credentials.md` for root vs. direct-spec modes, backend details, and the results-dir default. ## Container execution `tao-core` runs TAO containers through Pyxis/Enroot: 1. Stage compact JSON files for specs, environment, and cloud metadata under `/specs`, `/env`, and `/meta`. 2. Convert the Docker image to a cached SQSH image **before** the GPU job, with `srun -n1 -p enroot import`. This is a one-time cost per image, not an optional optimization — see *Acquire the image off the GPU allocation* below. 3. Write an sbatch script under `/sbatch/job_.sbatch`. 4. Submit `sbatch --export=ALL