--- name: tau-agent-performance description: Generate routine trailing-two-week provider/model latency and output-throughput CSV and SVG charts from content-free durable agent traces. --- # Routine agent performance charts Run the checked-in helper; do not reconstruct timing from captures or research the journal each time. It uses a native Cargo single-file Rust script, gnuplot, and Tau's validated `agent-performance-jsonl` trace projection. It never contacts a provider or running harness. The Nix runner supplies the nightly toolchain already pinned through Flakebox/Fenix in `flake.lock`, gnuplot, and a linker; it does not change the workspace's stable toolchain. First use can download/build the toolchain and script dependencies. ```sh cd "$(jj workspace root)" nix shell .#diagnostics -c tau-diagnostics-cargo \ .agents/skills/tau-agent-performance/chart_performance.rs \ --out tmp/agent-performance-last14days ``` The output directory must be new. Keep `performance.csv`, `latency.svg`, `throughput.svg`, their PNG previews and `.gnuplot` programs, `summary.txt`, and `README.md` together; do not commit them. Re-render with `gnuplot latency.gnuplot` and `gnuplot throughput.gnuplot` from the artifact directory. Programs contain only the same aggregate chart evidence, not traces. The helper prints the summary and artifact path. It creates an owner-only directory, but the artifacts still reveal aggregate activity and model names. Inspect them before sharing. ## Presentation defaults Deliver readable **PNG previews of both charts**, with the CSV, styled SVGs, and coverage summary available alongside them. The helper produces styled SVGs and 1600-pixel-wide PNGs using gnuplot, applying the defaults below. It recognizes family names only from the recorded model strings; no alias mapping is inferred. It builds one deterministic color map for the report, reserves the fixed palette, and resolves new-family hue/version-shade collisions over the present inventory. The saved programs retain the exact color bindings for reproducibility. For presentation tweaks, edit the saved private `.gnuplot` programs, preserving measurements and bucket gaps, and retain them for reproducibility. * Build one exact-model color map and reuse it for latency and throughput, across all accounts. Give model families stable hues and versions distinguishable shades: **Astra red, Sol green, Luna blue, Terra yellow**. Starting palette: Astra 6 `#dc2626`; Sol 5.6 `#15803d`, Sol 6 `#65a30d`, Sol 6.1 `#166534`; Luna 5.6 `#1e40af`, Luna 6 `#0284c7`; Terra 5.6 `#a16207` (dark golden yellow for readability on white); Grok `#be185d`; Qwen `#7c3aed`. Bind these colors to the actual recorded model strings. For new versions, extend the family's shades; for new families, choose a distinct hue. Family styling is presentation only: keep exact model identities and measurements separate, without guessing aliases or merging versions. * Use account marker shapes consistently: `chatgpt` circle, `chatgpt-fedi` square, `grok` diamond, `ren` triangle. Assign other recorded accounts distinct shapes. The CSV's `provider` is the recorded series key, not proof of account identity; use account labels only when that mapping is known, otherwise label the shape legend **Provider**. * Use two compact legends: **Model** with colored samples and exact model names, and **Account** (or **Provider**) with neutral marker shapes. Order model families **Astra → Sol → Terra → Luna**, then other families in stable alphabetical order; keep versions in stable ascending order within each family (compare numeric version components numerically). Reuse this order in both charts. Reserve enough space to keep legends and labels readable without covering data. * Prefer logarithmic y axes when positive values span a wide range. Label latency **seconds (log scale)** and throughput **tokens/wall-second (log scale)** when using log scales. Choose ticks and limits from the data, not fixed ranges that clip observations. Keep genuine zero rates visible using a linear scale or a clearly labeled separate zero indication; never turn missing values into zeros or substitute an epsilon on a log axis. * Show the UTC range and six-hour median bucketing. Preserve gaps, clipped boundary buckets, and today's current partial bucket. Include a short wall-time caveat and any material skipped-journal or sparse-coverage caveat in the delivery, using `summary.txt`. The helper renders SVG and PNG directly with the same gnuplot program. Inspect both PNGs at delivery size before sharing: confirm legible labels, unclipped legends, consistent colors/shapes, correct units, and honest zero/missing handling. Export the inspected PNGs through the artifact tool for inline previews; make the source artifacts available too. If no supported image viewer is available, report that visual inspection could not be completed rather than claiming it passed. **“Last two weeks” ends at the current moment, not the last midnight.** The default captures UTC now once before discovery/scanning, then selects the trailing fourteen days `[since, until)`. Today's current partial six-hour bucket is included. The companion `tau-qodq` quota/token workflow uses the same current-moment convention and clips token-rate denominators at the range boundaries. Neither workflow should round the endpoint down to midnight. For reproducible comparisons: ```sh nix shell .#diagnostics -c tau-diagnostics-cargo \ .agents/skills/tau-agent-performance/chart_performance.rs \ --agents-dir "$HOME/.local/state/tau/agents" \ --since 2026-09-17T13:25:00Z --until 2026-10-01T13:25:00Z \ --out tmp/agent-performance-fixed-range ``` Pass `--tau /path/to/tau` when the installed executable lacks current trace support. The default root is `$XDG_STATE_HOME/tau/agents`, falling back to `$HOME/.local/state/tau/agents`. Every immediate directory with `events.cbor` is scanned once, root-only, through a finite validated snapshot. There is no time index: a bounded output range does not reduce journal bytes scanned. Large histories can take several minutes. `--timeout` limits each individual trace command (default 120 seconds); it does not limit the full scan. The maximum range is 366 days. The helper retains selected scalar samples for medians and deduplication identities for the current journal, so memory scales with selected completed prompts plus prompt identities in the largest scanned journal, not raw trace bytes. Temporary traces are anonymous private files and removed automatically. Cargo builds and script lockfiles live in `$XDG_CACHE_HOME/tau-diagnostics/target` (default `$HOME/.cache/tau-diagnostics/target`), outside the source tree. Direct dependencies are exact-version pinned in the embedded manifest; Cargo keeps resolved transitive versions in its cached script lockfile, not in the repository. This is a pinned toolchain, not a fully vendored/offline script build. From the repository root, the executable `.rs` shebang invokes the same Nix runner. Do not use the unrelated third-party `cargo-script` or `rust-script` tools. ## What the charts mean Each line identifies the **exact recorded provider/model**, including version. The first slash of the canonical model separates provider from model; aliases such as sol/astra/terra/luna are not guessed or merged. Both charts use ordinary completed inference prompts in UTC-aligned six-hour buckets, assigned by the accepted terminal timestamp. The CSV reports completed, latency, and rate sample counts separately for every bucket/model: * **Latency:** median prompt-materialization to accepted response terminal journal-wall interval, in seconds. * **Throughput:** median of each prompt's `response_received_tokens / prompt-to-terminal elapsed seconds`. This is **not pure decoder speed**. Prompt preparation, provider work, retries, and harness overhead can be in the denominator. Retries are not separate chart samples. These are rough aggregate comparisons, not a wire profiler. First-output timing and separate harness overhead are unavailable in this durable projection. Standalone compaction is excluded. Missing terminals, missing timestamps, decreasing clocks, and missing usage are not zero measurements. Present zero tokens with positive elapsed time is a real zero-rate sample; zero elapsed time cannot supply a rate. Points sit at each clipped bucket's midpoint. Only adjacent buckets with that metric connect; missing buckets are gaps, not inactivity. Partial boundary buckets use only selected prompt samples; these per-prompt medians do not divide by bucket duration. No observations are fabricated at “now.” Check `summary.txt` before comparing providers: it reports scanned/successful, failed/unsupported/timed-out journals and selected/missing samples. Old schema journals can be unsupported; running agents can lack checkpoints or have a checkpoint behind their newest activity. Those journals are skipped explicitly, not treated as zero performance. Incomplete and missing-terminal-time counts cover all scanned history because their completion time is unavailable. Different workloads, reasoning effort, tool use, cache state, and retry rates can dominate model differences. A sparse median is only sparse evidence. ## Maintenance Reuse `docs/agent-trace.md` for the scalar schema and timing fidelity. Do not add provider captures, prompts, errors, or agent IDs to exported artifacts. The canonical trace projector owns typed correlation and journal validation. ```sh nix shell .#diagnostics -c tau-diagnostics-cargo test \ --manifest-path .agents/skills/tau-agent-performance/chart_performance.rs nix shell .#diagnostics -c tau-diagnostics-cargo test \ --manifest-path .agents/skills/tau-qodq/extract_quota.rs ```