# Apple Silicon M4/M5 Runbook This runbook makes the first Apple Silicon operational comparison repeatable: run the same benchmark packs on the local M5 Max machine and on an M4 Max Studio over SSH, pull the selected result directories back, and compare existing `run.jsonl` artifacts. The goal is a first-pass operational and performance comparison, not final proof of coding-agent quality. Do not commit generated benchmark artifacts from this workflow unless a later curated run-log entry explicitly calls for a small summary artifact. Raw responses under `results/*/raw/` remain local generated output. ## Prerequisites Both machines should have: - This repository checked out at the same commit or intentionally documented commits. - Dependencies installed with `uv sync`. - The same model installed and addressable by the runtime. - The same model tag, quantization, file, or adapter-visible model id. - The same runtime path where possible, such as both using `mlx_lm.server`, both using `llama-server` through the OpenAI-compatible adapter, or both using Ollama through the same adapter shape. - The same endpoint shape: OpenAI-compatible `/v1` base URL for `openai-chat`, or Ollama-native endpoint/defaults for `ollama-generate`. - Enough disk space for result directories, including `raw/`, `workspace/`, `patch/`, `task/`, and `verify/` artifacts for repo-task packs. The local M5 machine also needs SSH access to the M4 Studio: ```sh ssh 'uname -a' ``` Use placeholders for private details in notes and handoffs: ``, ``, ``, ``, and ``. ## Runtime Setup Start the same runtime/server shape on both machines before running packs. For OpenAI-compatible servers, use an endpoint base URL such as: ```text http://127.0.0.1:8080/v1 http://127.0.0.1:8081/v1 ``` For Ollama-native runs, the adapter can use its default local endpoint or an explicit endpoint if needed: ```sh uv run benchpack run smoke-chat --adapter ollama-generate --model --host-label m5-max-smoke ``` For `openai-chat` streaming packs, the default sends `stream_options.include_usage` so supporting endpoints can report token usage. If a local OpenAI-compatible server rejects that option, rerun the same command with: ```sh --openai-stream-usage omit ``` That compatibility mode preserves streamed output and TTFT, but usage-derived token counts and token-rate fields may stay null unless the server reports usage another way. ## Recommended Matrix Run these packs first: - `smoke-chat`: endpoint sanity only. Use it to prove the selected adapter, model id, endpoint shape, and basic scoring path work before spending time on larger packs. - `runtime-sweep`: the current performance comparison pack. Use it for TTFT, decode throughput, total throughput, prompt-token, cached-token, and gated prefill-parity comparison. - `desktop-django-wrap`: prompt-only coding-agent-shaped behavior. It exercises larger static prompts, fixture-backed prompt assembly, streaming, and regex scoring, but it does not copy, execute, or mutate a repository. - `patch-from-failure`: tiny repo-mutating, verifier-backed smoke benchmark using the current fenced unified-diff contract. It proves the narrow workspace, patch, task-log, and verifier path, not broad coding-agent quality. Use result labels that encode the host and pack. The default output directory is `results/-/`, so labels such as `m5-max-runtime` and `m4-max-runtime` produce distinguishable directories. Broad coding-agent conclusions still require production external harness execution, larger repo-task packs, and curated reporting around those runs. ## Default Qwen3.6 Targets These defaults are for explicit Qwen3.6 Apple Silicon comparison work and for continuity with the curated 2026-05-05 M4/M5 sweep. For current preferred model targets such as Gemma 4, consult `docs/model-targets.md` before launching a new campaign. When the goal is the current 30B-class Qwen comparison, run two model targets unless the run note says otherwise: - MoE: `Qwen/Qwen3.6-35B-A3B`. - Dense: `Qwen/Qwen3.6-27B`. For llama.cpp and Ollama, prefer one shared GGUF artifact per target so those two runtime paths use the same model file: - MoE GGUF: `unsloth/Qwen3.6-35B-A3B-GGUF`, `Qwen3.6-35B-A3B-UD-Q4_K_M.gguf`. - Dense GGUF: `unsloth/Qwen3.6-27B-GGUF`, `Qwen3.6-27B-Q4_K_M.gguf`. For the MLX path, use Apple-Silicon-native MLX conversions and document that the comparison is runtime-and-format, not bit-identical artifact parity: - MoE MLX: `majentik/Qwen3.6-35B-A3B-MLX-MXFP4`. - Dense MLX: `mlx-community/Qwen3.6-27B-4bit`. If a Qwen3.6 MLX conversion does not load through the selected MLX OpenAI-compatible server, record that runtime/model target as blocked. Do not silently substitute Qwen2.5, Qwen3-Coder, or another already-installed model when the run is meant to answer the Qwen3.6 question. ## Tmux Matrix Helper For metadata-backed matrix runs, prefer the narrow tmux helper over hand-typing four separate pack commands. The helper is only an operational wrapper around `benchpack run`: it does not discover runtimes, probe endpoints, change adapter requests, change result rows, alter pack semantics, run compare, or write reports. It keeps one tmux session open with deterministic windows so each pack's terminal output remains visible after the command finishes. Pack commands are gated to run sequentially, which avoids making the packs contend for the same local inference server. If one pack fails, later windows wake up, print that they were skipped because an earlier pack failed, and remain open for inspection. Create the metadata JSON first: ```sh mkdir -p metadata $EDITOR metadata/m5-llama-server.json ``` Then dry-run the local M5 commands. Keep placeholder values quoted until they are replaced, because shells treat angle brackets as redirection syntax: ```sh scripts/benchpack-tmux-matrix \ --dry-run \ --session-name 'bench-m5-llama-' \ --adapter openai-chat \ --model '' \ --endpoint '' \ --host-label-prefix 'm5-max-llama-' \ --run-metadata metadata/m5-llama-server.json ``` The default matrix is `smoke-chat`, `runtime-sweep`, `desktop-django-wrap`, and `patch-from-failure`. The helper maps those packs to host labels `-smoke`, `-runtime`, `-wrap`, and `-patch`. `--endpoint` may be omitted for adapters with useful defaults, such as a local Ollama-native run. `--openai-stream-usage include|omit` is optional; when it is omitted, the underlying `benchpack run` command keeps its current default. `--openai-api-key-env ` is also optional for authenticated OpenAI-compatible endpoints. The helper passes the environment variable name to `benchpack run` but does not read or print the token value. ## Known-good Qwen3.6 27B Strict-GGUF Helper Path The 2026-05-10 tri-host evidence validates one exact strict lane: `unsloth/Qwen3.6-27B-GGUF`, file `Qwen3.6-27B-Q4_K_M.gguf`, SHA256 `5ed60d0af4650a854b1755bd392f9aef4872643dc25a254bc68043fa638392a0`, alias `qwen36-27b-q4km`, served by `llama-server --reasoning off` on loopback `http://127.0.0.1:18082/v1`. For repeats of that lane, use the explicit preset after starting the matching server and preparing host-specific metadata: ```sh scripts/benchpack-tmux-matrix \ --dry-run \ --preset qwen36-27b-strict-gguf \ --session-name 'bench-m5-qwen36-27b-strict-' \ --adapter openai-chat \ --host-label-prefix 'm5-max-qwen36-27b-strict-' \ --run-metadata metadata/m5-qwen36-27b-strict.json ``` When omitted, the preset supplies `--model qwen36-27b-q4km` and `--endpoint http://127.0.0.1:18082/v1`. Explicit `--model` or `--endpoint` values override those defaults; the dry run shows the resolved command. If no positional packs or `--pack-set` are supplied, the helper still uses the default four-pack matrix: `smoke-chat`, `runtime-sweep`, `desktop-django-wrap`, and `patch-from-failure`. Combining the preset with `--pack-set` is supported when the campaign intentionally uses this strict-lane model and endpoint with a non-default pack set. This path does not launch `llama-server`, infer GGUF paths, create metadata, SSH to another host, pull results, or run reports. It also does not add `endpoint-python-correctness` to the four-pack helper path. Run that pack only as a separate explicit positional pack when the campaign calls for endpoint correctness evidence. Do not generalize this preset to Ollama, MLX, public API, or external-agent lanes. For optional exploratory repo-task evidence, dry-run a separate coding-task matrix instead of changing the default four-pack workflow. The existing `coding-tasks` set stays on the default fenced-patch repo-task harness: ```sh scripts/benchpack-tmux-matrix \ --dry-run \ --pack-set coding-tasks \ --session-name 'bench-m5-coding-tasks-' \ --adapter openai-chat \ --model '' \ --endpoint '' \ --host-label-prefix 'm5-max-coding-tasks-' \ --run-metadata metadata/m5-llama-server.json ``` `--pack-set coding-tasks` expands to `patch-from-failure`, `python-regression-fix`, and `django-dashboard-regression-fix`, in that order. Do not combine `--pack-set` with positional custom packs; the helper rejects that mix before generating commands. For external-agent evidence over the same three workloads, use the separate external-agent pack set and configure a real non-interactive agent command with `BENCHPACK_EXTERNAL_AGENT_ARGV`: ```sh BENCHPACK_EXTERNAL_AGENT_ARGV='["/path/to/agent"]' \ scripts/benchpack-tmux-matrix \ --dry-run \ --pack-set coding-tasks-external-agent \ --session-name 'bench-m5-coding-agent-' \ --adapter openai-chat \ --model '' \ --endpoint '' \ --host-label-prefix 'm5-max-coding-agent-' \ --run-metadata metadata/m5-llama-server.json ``` `--pack-set coding-tasks-external-agent` expands to `patch-from-failure-external-agent`, `python-regression-fix-external-agent`, and `django-dashboard-regression-fix-external-agent`, in that order. Those variants use the same fixtures and deterministic verifiers as the fenced-patch packs, but their prompts tell the external agent to edit the prepared workspace directly and select `harness = { id = "external-agent", timeout_s = 900 }`. The deterministic example harnesses under `examples/external-agent/` are for contract checks only; do not use them as live coding-agent evidence. For local Codex OSS/Ollama evidence, the source-controlled wrapper shape is: ```sh BENCHPACK_EXTERNAL_AGENT_ARGV="[\"python3\",\"$PWD/examples/external-agent/codex-oss-agent.py\",\"--codex-model\",\"qwen3-coder:latest\",\"--local-provider\",\"ollama\"]" ``` In launch mode, the helper requires `BENCHPACK_EXTERNAL_AGENT_ARGV` for this pack set and injects it into the tmux windows. Dry-run output names the requirement but does not print the value. Use an absolute wrapper path because the harness launches the external process from the prepared repo-task workspace. After inspecting the dry run, launch the tmux session: ```sh scripts/benchpack-tmux-matrix \ --session-name 'bench-m5-llama-' \ --adapter openai-chat \ --model '' \ --endpoint '' \ --host-label-prefix 'm5-max-llama-' \ --run-metadata metadata/m5-llama-server.json ``` The helper is conservative about result-directory collisions: it does not pass `--force` unless the command includes `--force`. Use `--force` only when the existing result directories are intentionally disposable. Launch mode also checks that the metadata file exists before it creates tmux windows; dry-run mode does not require the placeholder file to exist. Run the same helper on the remote M4 host through SSH after creating a remote metadata file in the remote repo: ```sh ssh ' set -eu cd uv sync mkdir -p metadata # Create metadata/m4-llama-server.json before launching the run. scripts/benchpack-tmux-matrix \ --dry-run \ --session-name "bench-m4-llama-" \ --adapter openai-chat \ --model "" \ --endpoint "" \ --host-label-prefix "m4-max-llama-" \ --run-metadata metadata/m4-llama-server.json ' ``` Then run the same SSH block without `--dry-run`. If the remote endpoint is bound to loopback, `` should be a loopback URL from the M4 Studio's point of view, for example `http://127.0.0.1:8080/v1`. After the local and remote tmux sessions finish, use the result pullback section below and generate comparison notes with `benchpack report` from the resulting directories. ## Local M5 Run The manual commands below are equivalent to the helper's default matrix and are useful when tmux is unavailable or when a single pack needs to be rerun. From the repo on the local M5 machine: ```sh uv sync uv run benchpack run smoke-chat \ --adapter openai-chat \ --model \ --endpoint \ --host-label -smoke \ --force uv run benchpack run runtime-sweep \ --adapter openai-chat \ --model \ --endpoint \ --host-label -runtime \ --force uv run benchpack run desktop-django-wrap \ --adapter openai-chat \ --model \ --endpoint \ --host-label -wrap \ --force uv run benchpack run patch-from-failure \ --adapter openai-chat \ --model \ --endpoint \ --host-label -patch \ --force ``` Example label choices: ```text = m5-max results/-m5-max-runtime/ results/-m5-max-wrap/ ``` If the OpenAI-compatible server rejects streaming usage options, add `--openai-stream-usage omit` to the streaming packs that need it, especially `runtime-sweep` and `desktop-django-wrap`. Before running packs that should be interpretable later, write a small metadata JSON file on each host and pass it with `--run-metadata`. Use placeholders in shared docs or handoffs when model paths or host details are private: ```json { "runtime": { "name": "llama-server", "version": "9010", "command": "llama-server --model --host 127.0.0.1 --port 8081 --ctx-size 4096 --gpu-layers auto", "options": { "ctx_size": 4096, "gpu_layers": "auto", "openai_stream_usage": "include" } }, "model": { "id": "", "source": "local-gguf", "quantization": "Q4_K_M", "sha256": "" }, "operating_conditions": { "power": "not captured", "thermal": "not captured", "background_load": "no intentional throttling setup" }, "notes": "optional short run note" } ``` Then add the flag to each matching run: ```sh --run-metadata metadata/.json ``` ## SSH M4 Run Run the same commands on the M4 Studio through SSH. Keep the remote repo path as a placeholder in docs and handoffs: ```sh ssh ' set -eu cd uv sync uv run benchpack run smoke-chat \ --adapter openai-chat \ --model \ --endpoint \ --host-label m4-max-smoke \ --force uv run benchpack run runtime-sweep \ --adapter openai-chat \ --model \ --endpoint \ --host-label m4-max-runtime \ --force uv run benchpack run desktop-django-wrap \ --adapter openai-chat \ --model \ --endpoint \ --host-label m4-max-wrap \ --force uv run benchpack run patch-from-failure \ --adapter openai-chat \ --model \ --endpoint \ --host-label m4-max-patch \ --force ' ``` If the remote endpoint is bound to the M4 Studio loopback interface, the `` value in the SSH command should usually be a loopback URL from the remote machine's point of view, for example `http://127.0.0.1:8080/v1`. For Ollama-native comparisons, keep the same host-label pattern and switch the adapter shape consistently on both machines: ```sh uv run benchpack run runtime-sweep \ --adapter ollama-generate \ --model \ --host-label m4-max-runtime \ --force ``` ## Result Pullback After the remote run, pull back only the result directories needed for the comparison. For compare-only work, `run.jsonl` is the required file; pulling `summary.md`, `hardware.json`, and `run-metadata.json` alongside it keeps the directory inspectable without copying large generated payloads. ```sh mkdir -p results/-m4-max-smoke mkdir -p results/-m4-max-runtime mkdir -p results/-m4-max-wrap rsync -a \ --include '/run.jsonl' \ --include '/summary.md' \ --include '/hardware.json' \ --include '/run-metadata.json' \ --exclude '*' \ :/results/-m4-max-smoke/ \ results/-m4-max-smoke/ rsync -a \ --include '/run.jsonl' \ --include '/summary.md' \ --include '/hardware.json' \ --include '/run-metadata.json' \ --exclude '*' \ :/results/-m4-max-runtime/ \ results/-m4-max-runtime/ rsync -a \ --include '/run.jsonl' \ --include '/summary.md' \ --include '/hardware.json' \ --include '/run-metadata.json' \ --exclude '*' \ :/results/-m4-max-wrap/ \ results/-m4-max-wrap/ ``` `scp` is also acceptable for the small compare files. Omit `run-metadata.json` from the command when that optional artifact was not captured: ```sh mkdir -p results/-m4-max-patch scp \ :/results/-m4-max-patch/run.jsonl \ :/results/-m4-max-patch/summary.md \ :/results/-m4-max-patch/hardware.json \ :/results/-m4-max-patch/run-metadata.json \ results/-m4-max-patch/ ``` Do not pull or add `raw/`, `workspace/`, `patch/`, `task/`, `verify/`, or other generated payloads for this slice. If a later run-log entry needs durable evidence, curate only the small artifacts called out in `docs/run-log.md`. ## Compare Workflow `benchpack compare` reads existing result directories that contain `run.jsonl` and writes only a textual comparison to stdout. Compare matching packs with matching result labels: ```sh uv run benchpack compare \ results/-m5-max-smoke \ results/-m4-max-smoke uv run benchpack compare \ results/-m5-max-runtime \ results/-m4-max-runtime uv run benchpack compare \ results/-m5-max-wrap \ results/-m4-max-wrap uv run benchpack compare \ results/-m5-max-patch \ results/-m4-max-patch ``` After checking individual comparisons, use `benchpack report` as the preferred way to assemble pasteable comparison notes. It reads the same `run.jsonl` rows as compare, includes optional `hardware.json` host identity, includes optional `run-metadata.json` runtime/model/operating details, counts scoring passes/failures, and reuses compare's median, warning, cache-row, and `prefill parity` logic: ```sh uv run benchpack report \ results/-m5-max-runtime \ results/-m4-max-runtime uv run benchpack report \ results/-m5-max-smoke \ results/-m4-max-smoke \ results/-m5-max-runtime \ results/-m4-max-runtime \ results/-m5-max-wrap \ results/-m4-max-wrap \ results/-m5-max-patch \ results/-m4-max-patch ``` For repeated report assembly, keep the result grouping in a source report-set manifest instead of pasting the full positional list every time: ```toml version = 1 result_dirs = [ "results/-m5-max-smoke", "results/-m4-max-smoke", "results/-m5-max-runtime", "results/-m4-max-runtime", "results/-m5-max-wrap", "results/-m4-max-wrap", "results/-m5-max-patch", "results/-m4-max-patch", ] ``` Then render the same Markdown report with: ```sh uv run benchpack report --set reports/.toml ``` Relative `result_dirs` entries resolve relative to the manifest file's parent directory. The manifest only names existing result directories; it does not run benchmarks, start tmux, contact the M4 host, pull files back, write report output, or mutate result directories. `benchpack compare` reports both emitted `WARNING:` lines and a table `prefill parity` column. Treat both as part of result interpretation, but keep their meanings separate. `prefill parity` column statuses: - `missing-case` means at least one compared result directory has no rows for that case. - `prompt-missing` means at least one non-empty case/run group lacks complete numeric `tokens.prompt` metadata. - `prompt-diff` means complete prompt-token medians differ, so cache and prefill conclusions are not comparable for that case. - `cache-missing` means prompt parity holds, but at least one side did not report complete `tokens.cached_prompt` metadata. - `cache-diff` means prompt metadata matches, but cached prompt-token medians differ. - `comparable` means every compared run has rows, complete numeric prompt/cache metadata, matching prompt-token medians, and matching cached prompt-token medians for that case. - `prefill_tps med` is rendered only when the case-level `prefill parity` status is `comparable`; all other statuses render `—`. Emitted `WARNING:` lines: - Different pack ids or versions mean the inputs are not reliable cross-pack comparisons. - Prompt-token median differences mean cache parity is not comparable across different prompts. - Incomplete cache metadata means some measured rows lack numeric `tokens.cached_prompt`. - Cached prompt-token median differences mean prefill speed should not be compared. Compare parity statuses and warnings are derived from normalized `run.jsonl` fields only. Compare does not infer runtime version, server command, quantization, model checksum, context size, GPU layer/batch/cache options, power state, thermal state, or background load from raw files, timing fields, or endpoint behavior. Capture those details with `benchpack run --run-metadata` whenever the comparison should be interpretable later. ## Hardware Metadata Check Before interpreting M4/M5 results, inspect each pulled `hardware.json`. For Apple Silicon comparisons it should identify the host through `chip`, `hardware_model`, `hardware_model_name`, `hardware_model_identifier`, `ram_mb`, `os`, and `gpus` when macOS reports those values. These fields distinguish host class, for example an M5 Max MacBook Pro from an M4 Max Mac Studio. They do not prove runtime parity. Use `--run-metadata` for runtime version, server command, model id, quantization, model checksum, context size, power mode, thermal state, and cache settings when a result is meant to be interpreted later. ## Comparison Report Checklist Before treating M4/M5 output as comparable, create a short report or run note that separates host metadata from `hardware.json`, user-supplied runtime metadata from `run-metadata.json`, and interpretation notes. Start with `benchpack report` output for the relevant result directories. The report output should replace manual copying of medians from multiple `benchpack compare` invocations and will include `run-metadata.json` when it is present. It still does not infer runtime versions, server commands, checksums, power state, thermal state, or background load; missing metadata must be called out as a comparability gap. Hardware identity from each result directory's `hardware.json`: - `chip` - `hardware_model` - `hardware_model_name` - `hardware_model_identifier` - `ram_mb` - `os` - `gpus` Runtime and operating metadata from `run-metadata.json`: - Runtime/server version and exact server command used on each host. - Adapter and endpoint shape, such as `openai-chat` with an OpenAI-compatible `/v1` base URL or `ollama-generate` with the Ollama-native path. - Model id, tag, or file path as recorded in private notes, plus quantization and checksum when practical. Use `` in shared docs or handoffs. - Context size, GPU layer settings, batch settings, prompt-cache or KV-cache settings, and the `--openai-stream-usage` mode used for streaming OpenAI-compatible packs. - Power mode, thermal state before the run, and meaningful background load. - Exact result directories compared. - `benchpack compare` warnings and `prefill parity` status by pack/case. - Per-pack interpretation: endpoint sanity for `smoke-chat`, performance for `runtime-sweep`, prompt-only behavior for `desktop-django-wrap`, and tiny repo-task smoke coverage for `patch-from-failure`. Compact report skeleton: ```text # Apple Silicon M4/M5 comparison report - Scope: - Repo commit on M5: - Repo commit on M4: - Packs compared: smoke-chat, runtime-sweep, desktop-django-wrap, patch-from-failure - Result directories: - M5 smoke/runtime/wrap/patch: results/-m5-max-... - M4 smoke/runtime/wrap/patch: results/-m4-max-... Hardware identity from hardware.json: - M5: chip=; hardware_model=; hardware_model_name=; hardware_model_identifier=; ram_mb=; os=; gpus= - M4: chip=; hardware_model=; hardware_model_name=; hardware_model_identifier=; ram_mb=; os=; gpus= Runtime and model notes: - M5 runtime/server: ; command: - M4 runtime/server: ; command: - Adapter/endpoint shape: against - Model: ; quantization=; checksum= - Runtime options: context=; gpu_layers=; batch=; cache=; openai_stream_usage= Operating conditions: - M5 power/thermal/background load: - M4 power/thermal/background load: Compare interpretation: - smoke-chat: endpoint sanity result only; do not use for performance claims. - runtime-sweep: note wall_s, TTFT, decode/total TPS, cache rows, warnings, and prefill parity status for short/medium/long before making speed claims. - desktop-django-wrap: prompt-only coding-agent-shaped behavior and streaming metrics; it does not mutate a repo or prove a wrap task. - patch-from-failure: tiny fenced-patch repo-task smoke; verifier pass/fail is useful for this fixture only and is not broad coding-agent quality evidence. Conclusion: - Comparable claims: - Exploratory or invalid claims: ``` ## Fairness Checklist Before interpreting M4-vs-M5 numbers, record or align: - Same model id, model file, or model tag. - Same quantization, such as the same MLX quant or GGUF quant. - Same runtime/server and version where possible. - Same adapter path, such as `openai-chat` on both machines or `ollama-generate` on both machines. - Same endpoint options, context size, GPU layer settings, batch settings, and prompt-cache settings. - Same `--openai-stream-usage` mode for OpenAI-compatible streaming packs. - Similar power mode and no intentional low-power throttling on one side only. - Thermal state before the run, especially whether either machine is already heat-soaked. - Background load, including other inference servers, indexing, builds, or backups. - Runtime version output and model checksum when practical. Treat M4-vs-M5 conclusions as invalid or exploratory when these items are not aligned or documented. ## Interpretation Boundaries `runtime-sweep` is the pack to use for performance comparison now. Its output is still only as fair as the model/runtime/cache alignment above. `desktop-django-wrap` is prompt-only. It is useful for coding-agent-shaped prompt behavior and streaming metrics, but it does not mutate a repository or prove a real wrap task. `patch-from-failure` is a tiny repo-task smoke benchmark. It exercises the current fenced patch, workspace, patch artifact, task log, and verifier path, but it is not enough for broad coding-agent conclusions. `python-regression-fix` is an optional next repo-task pack for the same fenced patch path. It is more realistic than `patch-from-failure`, but it is not part of the default four-pack M4/M5 matrix in this runbook. `django-dashboard-regression-fix` is an optional stronger multi-file repo-task pack for the same fenced patch path. Use `--pack-set coding-tasks` to run the current repo-task pack set as a separate exploratory matrix before treating its results as campaign evidence. Larger coding-agent claims should wait for production external harness support, larger repo-task packs, and curated reporting around those runs. ## Troubleshooting Endpoint smoke failure: : Run `smoke-chat` first and confirm the server is reachable from the machine executing `benchpack`. For SSH runs, `127.0.0.1` means the remote M4 Studio, not the local M5 machine. Server rejects `stream_options.include_usage`: : Add `--openai-stream-usage omit` to OpenAI-compatible streaming commands and note that usage-derived token counts and token-rate fields may be null. SSH quoting or path issues: : Keep `` free of shell-specific shortcuts in copied runbooks. If quoting becomes fragile, SSH into `` interactively and run the same commands directly from ``. Missing result directories: : Check the command's printed output path. `--host-label` controls only the default output directory name; `--out` overrides it entirely. A repeated label on the same date requires `--force` or a unique `--out`. Compare warnings: : `benchpack compare` uses existing `run.jsonl` rows only. Prompt, cache, and prefill warnings mean the compared rows are not equivalent enough for the affected conclusion, even when wall time or decode throughput numbers are still visible.