--- name: pytorch-profile-analysis description: "Analyze single-file PyTorch/Kineto Chrome trace `.json(.gz)` files using the VeloQ CLI. Use for CPU/CUDA/kernel correlation, ProfilerStep/annotation slicing, memory/shape grouping, and single-trace NCCL evidence." --- # PyTorch Profile Analysis Use `veloq pytorch` for PyTorch/Kineto Chrome traces: ```bash veloq pytorch summary T veloq pytorch search T --type kernel --name-regex 'nccl|gemm' --limit 20 veloq pytorch inspect T kernel:91 veloq pytorch correlate T kernel:91 veloq pytorch slices T --aggregate --group-by step veloq pytorch collectives T ``` This skill requires the VeloQ CLI on `PATH`. If `veloq` is missing, install it before analysis. ## Tool Boundary Use `veloq pytorch` verbs as the analysis interface. Do not query `.veloq/pytorch/` sidecars, generated Parquet files, or raw Kineto trace tables directly with DuckDB, PyArrow, pandas, or ad hoc SQL unless the user explicitly asks for raw-trace exploration or you are developing VeloQ itself. `veloq pytorch prep T` only builds/checks sidecars. After prep, continue with `summary`, `search`, `inspect`, `stats`, `correlate`, `timeline`, `slices`, or `collectives`. ## Inputs - Explicit `veloq pytorch` commands accept one Chrome trace named `.json` or `.json.gz`. - Automatic source detection only claims `.pt.trace.json` and `.pt.trace.json.gz`; explicitly select `pytorch` for other JSON filenames. - Directory inputs are not supported in PyTorch v0. Ask the user to choose one trace file if they point at a directory. ## Row IDs PyTorch row ids use `:`, where the stable index is derived from the original `traceEvents` order after non-event flow markers are skipped. Do not use Kineto `Ev Idx` as a stable key. Use `veloq pytorch schema ` for the authoritative response field inventory; do not infer the public contract from raw Kineto fields. Common prefixes: | Type | Row id prefix | | ---------- | -------------- | | CPU op | `cpu_op:N` | | Annotation | `annotation:N` | | Step | `step:N` | | Runtime | `runtime:N` | | Driver | `driver:N` | | Kernel | `kernel:N` | | Memcpy | `memcpy:N` | | Memset | `memset:N` | | Memory | `memory:N` | | Python | `python:N` | | Comm | `comm:N` | ## Workflow 1. Inventory first: ```bash veloq pytorch summary T ``` Read `data.auxiliary.capabilities` before choosing a path. 2. Find events: ```bash veloq pytorch search T --type cpu-op --name '*aten::*' --limit 20 veloq pytorch search T --type kernel --is-comm --limit 20 ``` 3. Drill into one event: ```bash veloq pytorch inspect T ROW_ID ``` Inspect returns raw args, typed args, parent/children, enclosing step, and link metadata. 4. Answer launch-cause questions: ```bash veloq pytorch correlate T kernel:91 ``` Read `data.rows[0].events[]` for the CPU op, annotation/step, runtime/driver, and GPU activity chain. 5. Attribute CPU overhead to Python context when captured: Traces exported from `torch.profiler.profile(..., with_stack=True)` include `python_function` events. `inspect` returns `python_context` / `python_stack`, and `stats` can group CPU work by `python-context` or `python-path`: ```bash veloq pytorch stats T --type cpu-op --group-by python-path,name --limit 20 veloq pytorch inspect T cpu_op:42 ``` 6. Slice ProfilerStep/user annotation ranges: `slices --from/--to` selects ranges that overlap the time window and clips attributed GPU/comm time to that window. Slice row `start_ns` and `duration_ns` remain the original trace range so the row id still points to the inspectable event. 7. For communication questions, stay within one trace file: ```bash veloq pytorch stats T --type comm --group-by comm-kind,rank veloq pytorch search T --type kernel --is-comm --limit 20 ``` `collectives` groups single-trace communication evidence and reports linked CPU/NCCL row ids. If a trace file contains multiple rank values, rank-scoped commands (`search`, `stats`, `timeline`, `slices`, and `collectives`) require `--rank ` or `--all-ranks`. `inspect` and `correlate` operate on explicit row ids and are not rank-scope gated. Device ids are rank-local and stream ids are device-local: filter a stream with `--rank --device --stream `, or compare lanes with `--group-by rank,device,stream`. VeloQ does not compute cross-rank skew in PyTorch v0: ```bash veloq pytorch collectives T ``` ## Event Types `--type` accepts `cpu-op`, `annotation`, `step`, `runtime`, `driver`, `kernel`, `memcpy`, `memset`, `memory`, `python`, `comm`, or `all`. `comm` is a communication-related set. Use `--type kernel --is-comm` to focus on NCCL kernels. ## Limits PyTorch support is experimental (`source.version = "v0"`). Classification is based on Kineto category/name/arg conventions and may need extension for profiler variants not yet represented by tests. Treat documented fields, schema targets, row ids/keys, command ids, and output modes as the versioned source contract even while the source remains `v0`.