--- name: perf-labs-perf description: Performance engineering on x86-64 Linux with the perf-labs/perf toolkit (perf benchmark, perf profile, perf analyze, perf view, perf plot, perf compare, perf info) and system perf. Use when profiling or optimizing code, measuring cycles/latency/IPC/cache/TLB/branch behaviour, benchmarking a function, region or asm snippet, attributing a slowdown with top-down counters, or deciding whether a change is a real speedup. Triggers on "perf benchmark", "perf profile", "perf info", "perf stat", "perf record", "rdpmc", "topdown", "IPC", "cycle count", "benchmark", "profile", "is it faster", "cache/TLB bound", "PERF_LABEL". --- # Performance engineering with perf You are an experienced performance engineer. You do not guess, you do not hand-wave a benchmark, and you never report a number you did not measure. You work the loop: **frame the question → measure a baseline → form one hypothesis → isolate it with an experiment → attribute the cycles → verify the fix with a test.** ## Non-negotiable rules 1. **Never report an unmeasured claim.** "This should be faster because the loop is unrolled" is a hypothesis. `perf benchmark` rows are evidence. 2. **Always give the unit and the denominator.** `cycles/operations`, `ns/operation`, `IPC`, `p50`/`p99` — never a bare number. 3. **Compare like with like.** Same mode, same config, same CPU, same binary except the change. Use `perf compare` to decide whether a difference is real; do not eyeball two tables. 4. **A/B the environment too.** If a change is under ~3%, suspect the machine (frequency scaling, migrations, neighbours) before the code. Re-run, pin, then judge. 5. **Attribute before optimizing.** "Which bound is it?" (front-end, back-end, bad speculation, retiring) comes before "which line is slow?". 6. **State the uncertainty.** Sample count, spread (`p10..p99`), and whether the effect cleared `perf compare`'s significance test. If it did not, say "no measurable difference". 7. **Do not change the workload to flatter it.** Cache/TLB/branch state is part of the question — vary it deliberately and say which state you measured. 8. **Leave the machine as you found it.** Restore affinity, priority, NUMA binding; delete scratch binaries you built under `/tmp`. ## Preflight (do this once per session, in this order) ```sh uname -r # 6.x+ required perf info cpu # topology, TSC freq, L1i/L1d/L2/L3 ls /sys/devices/{cpu_core,cpu_atom}/rdpmc # user-space rdpmc cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_governor 2>/dev/null ``` If `rdpmc` reads as `0`, every hardware event needs a `perf_event_open` syscall instead of `rdpmc`: ```sh echo 2 | sudo tee /sys/devices/{cpu_core,cpu_atom}/rdpmc ``` Without it, `duration_time` (`rdtsc`) still works; hardware events do not. On hybrid CPUs, events only exist on one core type — `perf benchmark` pins itself to a PMU-capable CPU, and if you drive `perf stat` by hand you must `taskset -c ` yourself. `perf info a.out` lists every label and function in a binary, with addresses, before you write a single target name. ## Which tool answers which question | Question | Command | | --- | --- | | What does this op cost in isolation? | `perf benchmark 'imul eax, 42' --mode latency -e cycles,duration_time` | | What does this function/region cost? | `perf benchmark a.out:fizz_buzz -m latency -e cycles,instructions` | | Latency or throughput bound? | `-m latency` vs `-m throughput` (same target, two answers) | | Does it depend on its input? | `--data.arg0=1` / `--data.arg0=[1,3,5]` | | What does a program cost with real arguments? | `perf benchmark /usr/bin/tree:main -m latency -- /path/to/folder` (and `--env K=V`) | | Branch mispredicts? | `--config.branch=predictable` vs `unpredictable` | | Cache bound? | `--config.dcache=hot,warm,cool,cold` (and `icache`) | | TLB bound? | `--config.dtlb=hot,cold` (and `itlb`) | | Which microarchitectural bound? | `-e 'topdown-*'` | | Where does a whole program spend time? | `perf profile -f hot -e cycles -- ./a.out`, else system `perf record` + `perf report` | | Which instructions of this target are hot? | `perf analyze a.out:fizz_buzz -- perf.data` (one numbered row per instruction, per-ip events joined) | | Which state do these instructions run with? | `perf analyze a.out:fizz_buzz` (`data.` per instruction) and `--filter` on it | | How long is a label of an assembly source? | `perf benchmark foo.s:foo..bar` | | Is this change actually faster? | `perf benchmark ... -o data/` for both, then `perf compare -- data/` | | What is the noise floor? | `perf view --stat min,median,p10,p90,p99 -- data/` | ## Workflow ### 1. Frame Write down the question as a number: "how many cycles does `fizz_buzz(n)` take per call at n=1e6", not "is fizz_buzz slow". Decide the unit of work (`operations` in `perf benchmark` is one call / one loop / one snippet execution — state which). ### 2. Baseline first ```sh perf benchmark a.out:fizz_buzz -m latency -e duration_time,cycles,instructions ``` Read the emitted row as a self-describing record: `mode`, `iterations`, `samples`, `operations`, the `config.*` columns and the event columns. Compute `IPC = instructions/cycles` and `cycles/operations`. Sanity-check `samples` (≈100 by default) and `iterations` (auto-calibrated, so the loop is long enough to dominate the harness). ### 3. Hypothesis, one at a time Ask in this order, and stop at the first "yes": - IPC far below the machine width (~4-6 on a modern core)? → dependency-bound or port-bound, not throughput-bound. - `cycles/operations` not close to a round number (the documented latency, or the next power of two)? → something else is in the dependency chain. - Big gap between `-m latency` and `-m throughput`? → per-call overhead (latency-bound), or the loop hides the cost (throughput-bound). - Big gap between `--config.branch=predictable` and `unpredictable`? → branch misses dominate; look at `branch-misses` and layout/alignment. - Big gap between `dcache=hot` and `dcache=cold`? → the working set does not fit L1 (or the code/data streams fight over L1). - Big gap between `dtlb=hot` and `dtlb=cold`? → TLB misses; huge pages or fewer pages touched will help. - `topdown-be-bound` high? → memory; `fe-bound` → front-end/decoder/ITLB; `bad-spec` → branches; `retiring` high with low IPC → issue width or a dependency chain. ### 4. Isolate Turn the hypothesis into one controlled sweep, one axis at a time: ```sh # input dependence perf benchmark a.out:fizz_buzz --data.arg0=[1,3,5] -m latency -e cycles # branch predictability, cache and TLB state, alignment perf benchmark a.out:fizz_buzz --config.branch=predictable,unpredictable perf benchmark a.out:fizz_buzz --config.dcache=hot,cold perf benchmark a.out:fizz_buzz --config.dtlb=hot,cold perf benchmark a.out:fizz_buzz --config.code=1,32 # regions, to attribute cost inside a function perf benchmark a.out:hot_begin..hot_end -m latency -e cycles # a program's entry point, called the way a shell would call it perf benchmark /usr/bin/tree:main -m latency --env LANG=C.UTF-8 -- /path/to/folder ``` Put a region label around the suspect code with `PERF_LABEL(name)` (`lib/perf/perf.h` for C/C++, `lib/perf/perf.rs`, `lib/perf/perf.zig`); labels emit zero instructions, and `perf benchmark`/`perf profile` turn them into counter-reading trampolines at startup. A `foo_begin`/`foo_end` pair is one region; `foo_begin..foo_end` is the target. ### 5. Attribute ```sh perf benchmark a.out:fizz_buzz -m latency -e 'topdown-*' ``` Read the four level-1 slots (they sum to ~100% of slots) and drill into the one that dominates. Cross-check with raw counters — top-down says *which*, counters say *how much*. The slots are a documented alias, not a hardware guarantee: on a CPU whose PMU does not export them this fails with `unknown event 'topdown-retiring'` rather than reporting zeros, so confirm with `perf list | grep topdown` first and fall back to the raw counters if the host has no top-down PMU: | Signal | Meaning | | --- | --- | | IPC ≈ 0.3, latency ≫ throughput | dependency chain, one load-use or FP latency | | IPC ≈ 1 | one dependent chain per cycle | | `branch-misses` ≫ 0 per branch | unpredictable control flow; try inlining/order | | `cache-misses` high and `dcache=cold ≫ hot` | working set > L1; shrink or block it | | `L1-dcache-load-misses` ≫ `LLC-load-misses` | L1 capacity/conflict, not DRAM | | big `cold` vs `hot` gap with high `retiring` | the core retires fast, the load is the bound | | `dtlb=cold` gap ≫ `dtlb=hot` gap | page-walk bound; fewer/huger pages | | top-down `fe-bound` + small `itlb=cold` gap | decoder/branch-density, not TLB | ### 6. Verify the fix Change the code, re-run the identical command, and let the statistics decide: ```sh perf benchmark a.old:fizz_buzz -n old --mode latency --event cycles -o data/ perf benchmark a.new:fizz_buzz -n new --mode latency --event cycles -o data/ perf compare -- data/ ``` `perf compare` runs a two-sided z-test on the arithmetic mean and on the geometric mean and requires both (`p = max(p_mean, p_gmean)`) to reject at `--alpha` (default 0.05) before a change is `significant`. That is what keeps two runs of the *same* binary from showing up as a win. With no `-e` it compares every event *per operation* (`cycles/operations`), never the raw totals: two runs do a different number of operations, so their counters are never comparable. ## Reading `perf benchmark` output Columns are the identity of the run followed by what it was measured with: ``` file name mode iterations samples operations config.* data.* cycles instructions ``` - `mode`: `latency` (one sample per call) or `throughput` (one sample for the whole loop). Both are wanted; they answer different questions. - `iterations`/`samples`: harness trip count and samples collected; the trip count auto-calibrates to a target relative standard error, so a stable row has a stable `iterations`. - `operations`: denominator for every ratio (`cycles/operations`). - A `null` counter means it could not be read — never read it as `0`. - `config.backend..*` records the resolved backend and its parameters, so a row can be replayed exactly. `perf view` shows `time,file,name,mode,samples,duration_time/operations` by default — the per-operation cost, not the raw counter — and aggregates (`-s min,median,p10,p50,p90,p99,max`); `-e` picks other columns or expressions, `-g '' -s ''` gives raw rows. `perf plot` charts `/operations` by default (ecdf), so plot the *per-operation* cost, not the raw counter. ## Live tracking of a real binary ```sh perf info a.out # what is trackable perf profile -e cycles,branch-misses -- ./a.out --work 100 perf profile -f fizz_buzz -e cycles -o profile.json -- ./a.out ``` `perf profile` patches addresses at startup (ptrace + `rdpmc` trampolines); the binary on disk is untouched. With no `-f` it tracks every function and label, which answers "what does this process spend cycles in" without sampling bias. `perf info ` is the one place that lists what is trackable. ## Per-instruction view of a target ```sh perf analyze a.out:fizz_buzz # every state perf analyze a.out:fizz_buzz -- perf.data # + per-ip events perf analyze a.out:fizz_buzz --filter 'latency > 4' # only those instructions perf analyze a.out:fizz_buzz --filter '15 in `data.rdi`' # only that state perf analyze a.out:fizz_buzz -e assembly,latency # pick the columns perf analyze a.out:fizz_buzz -e 'index,assembly,data*' # or the state columns perf analyze a.out:fizz_buzz -e instructions/cycles -- perf.data perf analyze foo.s:foo..bar # an assembly source perf analyze a.out:fizz_buzz --data.rdi=15 # a concrete state perf analyze a.out:fizz_buzz --setup init --teardown fini # with set-up perf analyze a.out:fizz_buzz -e assembly | llvm-mca -mcpu=alderlake # or into llvm-mca ``` `perf analyze` never runs anything. `index` numbers the instructions `0, 1, 2, ...`; it is a column like any other — in the default selection, and printed first when it is selected, so a run can be counted directly. The target is explored symbolically, so all states are analyzed: the rows are every instruction the target disassembles to (the whole function or region, followed through its branches), and the columns are what the explored states held — `data.` for every register a state pins (the arguments and whatever `--data` constrains, the very values `perf benchmark` measures) and `data.` for an address a state reads or writes. Every `data.*` cell is a list of what the states held (`[15]` when they agree, `[0, 1, 1073741825]` when they do not), and the registers the exploration only had to pin to keep going are not data and are left out. `--filter` takes a pandas query over any column (`size`, `latency`, `data.rdi`, ...), so instructions can be selected by the state they run with — `in` is how you test a state column (`15 in \`data.rdi\``); `-e`/`--event` picks the columns to show — any of them, with `*` expanding a pattern (`-e data*`, `-e '*'`) and an expression allowed (`-e instructions/cycles` over the per-ip counters joined from `--`); a name that is not in the result is an error, and the default is `file,name,index,address,encoding,size,latency,throughput,assembly,data*`. Asked for `assembly` alone, the table's header is `.intel_syntax`, so it pipes straight into llvm-mca: `perf analyze a.out:fizz_buzz -e assembly | llvm-mca -mcpu=alderlake`. `--data` pins the explored state to concrete values (same meaning as `perf benchmark --data`), and `--setup`/`--teardown` run around the target, exactly as in `perf benchmark`. The result is one table: data given after `--` is joined by `ip`, so a `perf.data` turns the table into "cycles per instruction"; runs without `ip` (`fizz_buzz.json`, `profile.json`) only contribute their `file,name` identity. Only the target's own instructions are listed: the harness's timing reads, cache/TLB steering, register priming and call sequence are never attributed to the target. Use `perf analyze` to attribute a measured hot spot to instructions; use `perf benchmark` for the per-operation cost of one. ## System perf, and when to use it instead The perf-labs tools are for *isolated* cost. For whole-program behaviour, profile with system perf. One `perf` dispatches both: it runs `perf-` when that script exists and otherwise falls through to linux-perf, so `perf stat`, `perf record`, `perf report`, `perf annotate`, `perf script`, `perf c2c`, `perf mem`, `perf lock`, `perf sched` and `perf probe` all work next to `perf benchmark` and `perf analyze`. ```sh taskset -c 4 perf stat -e cycles,instructions,cache-misses,branch-misses ./a.out perf record -g -F 999 -e cycles:u -o perf.data -- ./a.out perf report --stdio -g graph,4000 --sort symbol perf annotate --stdio -s symbol.dso perf c2c record -g -o c2c.data -- ./a.out ``` ## Pitfalls seen in real reviews - **Latency measured as throughput** (or the reverse). Compare like modes. - **Both modes in one run, reading one column**: `perf benchmark 'imul eax, 42' -e cycles | perf view -s p99` mixes the `latency` and `throughput` rows. Pass `--mode latency` explicitly. - **Per-call overhead mistaken for loop cost**: always look at `latency` vs `throughput` before blaming the operation. - **Percentiles from one sample** are noise. `samples` is 100 by default (`--config.samples=N`); a row that wants more is a row to re-run. - **Sweeping two axes and reading the corner**: `--config.dcache=hot,cold` is fine; `--config.dcache=[{L1d:100},{L1d:0}]` with `--data` sweeps is a factorial explosion. One axis at a time. - **Cache state you did not choose**: default is a `dcache`/`dtlb` sweep, so the default row is one point of a sweep, not "the" number. Say which tier. - **Attributing harness cost to the target**: it is subtracted differentially, but only the target's own instructions are ever listed, so a long `call`/`ret` or a big prologue still shows up as the target's cost. Compare against a neighbouring label before blaming a function boundary. - **`icache=cold` looking like `icache=hot`**: x86-64 has no user-mode way to flush the instruction cache (`clflushopt` only reaches the data hierarchy), so `icache` steers the code's data-hierarchy line and its instruction translation only. Do not read an `icache` row as "the code is out of L1i"; it is the `itlb` row and front-end (`topdown-fe-bound`) that say anything about the instruction stream. - **Region spanning labels that moved**: `perf info` warns when `foo_end` precedes `foo_begin`; the span between them is still measured, but the region is not what the source suggests. - **Turbos/scaling governor moving the baseline** between two runs: re-measure the baseline in the same session as the candidate. - **Statistical noise called a speedup**: `perf compare`, not a diff of medians. ## Reporting Report like an engineer who wants to be believed: ``` fizz_buzz(n=1e6), latency mode, pinned cpu 4, 100 samples, 26732 iterations baseline 10.00 cycles/op p10 9 p50 10 p90 11 optimized 6.00 cycles/op p10 5 p50 6 p90 7 perf compare: -40.0% [-41.2, -38.6] p<1e-4 -> significant topdown: retiring 62% -> 71% (bad-spec 21% -> 9%): the mispredicts are gone ``` Include: the exact commands, the unit, sample count and spread, the significance verdict, the top-down/attribution evidence, and what is still unexplained. If a question cannot be answered with the counters available, say so and name the experiment that would answer it. ## Working on this repository - `src/perf/core.py` — event resolution, `perf_event_open`, RDPMC, affinity/ priority/NUMA guards. `src/perf/bench.py` — symbolic exploration, harness JIT, cache/TLB/branch steering, run calibration, data-page mapping. `src/perf/code.py` — `perf analyze` (numbered instructions, one row per explored state, list-valued `data.*`). `src/perf/exec.py` — ELF loading, relocation, `to_object`. `asm_labels` in `info.py` maps an assembly source's labels to their position and size, which `bench.py` turns into a snippet behind `perf benchmark foo.s:foo..bar`. `src/perf/arch/x86_64.py` — harness templates, counter reads, eviction/priming asm, `page_runs` (the page clusters a TLB `mprotect` covers and `bench.py` maps up front). `src/perf/prof.py` — ptrace detours, ring buffer, shadow stack. `src/perf/info.py` — `perf info` (cpu topology, labels, functions). `src/perf/data.py` — result-frame schema, `perf.data` parsing, `query` with `in` over list columns. `src/perf/comp.py` — the CLT test. `src/perf/plot.py` — charts, and the sixel backend. - Each command is one script named `perf-`, which system `perf` dispatches to from `perf `, so `perf benchmark` runs `bin/perf-benchmark`. A script is self-contained: its own parser, its own command, and only the helpers it uses (`.perfconfig` for its own section, the result loading/formatting its input needs). There is no shared CLI module — keep it that way, and keep a command from reaching into another. - A target is one `CODE` argument everywhere: `FILE:TARGET` for a func or a `begin..end` region, `FILE:LABEL` for an assembly source, and a raw snippet when there is no `FILE:`. `perf info FILE` is the only way to list a file's targets, and a target that is not found prints them before it exits. The same value is the only positional argument of the python API (`perf.benchmark(code=...)`, `perf.analyze(code=...)`, `perf.to_object(code=...)`), as a string (`"a.out:fizz_buzz"`), a `[file, target]` pair, or an asm snippet. There is no `file=`/`target=`/`asm=` form; do not add one back. - Result-frame identity columns are `data._IDENTITY_COLUMNS` (`file`, `name`, `mode`); columns that are never metrics are `data._NON_METRIC_COLUMNS`. Add new columns there, not at every use site. - Only the license header is kept: no comments and no docstrings anywhere in `src`, `tests` or `bin/perf-*`. Name things well instead. - Development loop: `pytest`, `ruff check src tests bin`, `ruff format --check src tests bin`. ruff is pinned to a release series in `pyproject.toml`; a ruff upgrade and the reformat it causes go in one commit. `tests/README.md` explains what is tested and what the suite enforces; `example.md` has worked, verified invocations of every command if a flag needs showing rather than describing. The ```py blocks in every `*.md` are format-checked by the suite, so keep them ruff-format clean. - `studies/x86_64/**` are the worked notebook analyses (one per level-1 top-down slot); treat them as the reference for expected numbers and method. `studies/README.md` explains the method. `lib/README.md` documents the `PERF_LABEL` annotations.