--- name: pbs-hpc description: Prepare PBS HPC jobs, choose allocation and filesystem settings, submit and monitor jobs, and diagnose job output. Use for qsub, qstat, PBS scripts, allocations, and ChemGraph execution on Polaris, Aurora, or Crux; also Polaris and Aurora environments, MPI/GPU affinity, monitoring, and Aurora /soft access troubleshooting. compatibility: Submission requires PBS commands on an authorized submission host; compute execution requires an allocation and the site's application environment. license: Apache-2.0 metadata: authors: Murat Keceli maintainers: tdpham2 --- # PBS and HPC workflows ## Establish the environment Identify the target system, submission host, project/account, queue, walltime, node count, required filesystems, and application environment. Obtain missing values from the user or deployment configuration. Do not copy another user's allocation, endpoint, private path, or environment. Check whether `execute` runs on that host. A local shell on a laptop cannot run remote PBS commands merely because a skill or an HPC MCP server is attached. When the shell runs on the submission host, write the job files and submit with PBS commands directly. Otherwise use suitable attached HPC tools, or prepare the script and state where the user must submit it. Read the applicable site reference before choosing launch settings: [Polaris](references/polaris.md), [Aurora](references/aurora.md), or [Crux](references/crux.md). Consult the linked official guides for current queue limits and site policies. The Polaris reference includes login-node usage, modules, network proxy setup, queues, MPI/OpenMP examples, CPU/GPU affinity, MPS/MIG, and storage. Read it for Polaris environment setup as well as job submission. The Aurora reference covers hardware, queue selection, PALS, Intel GPU hierarchy and affinity, monitoring, and group-restricted `/soft` access. Read it for Aurora operations and environment troubleshooting as well as job submission. ## Prepare and track a job For ASE calculations, also read the `chemgraph` skill's [Python batch example](../chemgraph/references/ase-batch.md). Write the calculation script in the workspace and run it inside PBS; no MCP server or Parsl is needed. 1. Read [the PBS template](assets/job.pbs.template) and write a completed copy into the execution filesystem. Replace every placeholder, keep PBS directives before executable statements, and quote shell paths. Choose the launch command for the target site and application. 2. Check input visibility on compute nodes. Arrange staging first when the submitting host and workers do not share files. Validate the script syntax with `bash -n` when the execution environment provides Bash. 3. Submit once on the authorized submission host, following existing tool approvals. Save submission evidence and the complete scheduler ID in the run directory. 4. Inspect with `qstat -f JOB_ID`; distinguish queued, held, running, and terminal states. Explain scheduler comments rather than submitting duplicate jobs. 5. Inspect stdout/stderr and application result files. A job disappearing from the active queue does not prove success. Use available job history and exit status, and report missing evidence explicitly. Cancel only the requested job. Run this submission sequence from the fresh host run directory after completing and inspecting the files: ```bash set -euo pipefail bash -n job.pbs test ! -e job.id if ! (set -C; : > submission.started) 2>/dev/null; then echo "Submission already attempted; inspect PBS before retrying." >&2 exit 1 fi qsub job.pbs > job.id 2> qsub.stderr test -s job.id cat job.id ``` Keep the marker and stderr if `qsub` fails or returns no ID: acceptance can be uncertain. Inspect PBS by job name and run directory before any retry. A later agent session reads the existing `job.id`, checks that job with `qstat -f` (or `qstat -xf` for retained history), and inspects output files without resubmitting. An explicitly requested retry uses a fresh directory after resolving the prior job's state. Cancellation uses `qdel JOB_ID` only when requested. PBS command reference: [ALCF running jobs](https://docs.alcf.anl.gov/running-jobs/). ## ChemGraph execution boundaries ChemGraph's `hpc_configs` Parsl configurations for these systems use `LocalProvider` inside existing allocations; they do not acquire a PBS allocation automatically. Aurora and Crux require `PBS_NODEFILE`; Polaris has a local/testing fallback, which is not evidence of a valid allocation. Use `CHEMGRAPH_WORKER_INIT` or the configured Python environment for worker setup. Check the deployment's shared filesystem assumption: Globus Compute workers may not see files written by the submitting server. A Globus endpoint has its own execution configuration and is distinct from direct `qsub` usage.