# Prefill on the NPU (`--prefill-device`) Prefill and decode want opposite hardware. A prefill graph is hundreds or thousands of tokens wide, so a matrix engine runs it many times faster than the CPU cores. A decode graph is one token wide, and on a phone's unified memory the cost of crossing to an accelerator and back eats what it saves: measured on a Hexagon NPU, the decode token got slower (see "Why not decode" below). `--prefill-device D` gives each phase its hardware. Wide prefill graphs run on ggml device `D` (the Hexagon NPU is `HTP0`); decode stays on the CPU, exactly as without the flag. With `--moe-stream` it works on a model larger than RAM: the NPU never holds the model, only two layers of it at a time. ## Measured Qwen3.6-35B-A3B, Q4_0 gguf (20.8 GB) on a 12 GB phone with a Hexagon v81 NPU and UFS 4 storage, streamed (`--moe-stream --overlap --dense-weights ahwb --cache-mb 1500 -t 4`), ubatch 2048 on the NPU, 512 on the CPU. The Q4_K_M row is the same model in the quantisation the app's catalog ships, on the 2026-09-26 llama.cpp base, with 8 loaders: | prompt | CPU prefill | NPU prefill | | |---|---:|---:|---:| | 121 tokens | 9.95 s (12.2 tok/s) | 9.5 s (12.7 tok/s) | parity: flash bound | | 1418 tokens, prose | 63.8 s (22.2 tok/s) | 8.2 s (172 tok/s) | 7.8x | | 1921 tokens | 106.5 s (18.0 tok/s) | 11.2 s (171 tok/s) | 9.5x | | 1418 tokens, prose, **Q4_K_M** | 80.6 s (17.6 tok/s) | 10.0 s (141 tok/s) | 8.1x | Decode after the 1418-token prompt: 3.41 tok/s on the CPU-only run, 3.25 tok/s with the NPU prefill (both decode on the CPU; the difference is the memory the NPU's slots hold). Gemma 4 26B-A4B-it Q4_K_M (17 GB), same phone and settings but `--cache-mb 2000`, a 238-token prompt, a 2048-token context: | | prefill | decode | |---|---:|---:| | CPU | 16.2 s (14.7 tok/s) | 3.60 tok/s (8192-token context) | | NPU | 5.85 s (40.7 tok/s) | 3.30 tok/s | **Memory is the limit on a model with a large KV cache.** The device path holds the two expert slots (1.1 GB on this model), the device's compute buffers (0.8 GB at ubatch 2048) and the model state in the device's host buffer. On a 12 GB phone Gemma 4 at an 8192-token context and a 2000 MiB expert cache already leaves about 330 MB free on the CPU alone; with the device path on top it thrashes. At 2048 it fits. A smaller `--cache-mb` or context makes the room. **Short prompts are flash bound.** The table above was measured reading the whole expert set from flash once per graph (17.4 GB here), whatever the prompt length. Past roughly a thousand tokens that read hides behind the NPU's compute; under a few hundred it is the whole cost. The arena now reads only the experts a graph routes to (below), which on short prompts is about half of them. ## How it works ### Moving the weights per graph, not per model llama.cpp's scheduler runs each op on the backend that holds its weight, and it decides that again for every graph it builds. Where a weight lives is three public fields of its `ggml_tensor`: `buffer`, `data`, `extra`. So before a wide prefill graph the engine rebinds the layer weights onto the NPU's buffers, and afterwards back onto the CPU's. Nothing in llama.cpp is patched. The device joins the scheduler at load as the model's only non-CPU device, with no layer assigned to it, and `op_offload` is off, so the device runs nothing it was not handed. The one hazard is graph reuse: llama.cpp skips re-scheduling a graph shaped like the previous one. The rule that makes the rebind safe is about widths. The prompt goes in one ubatch per decode; pieces at least `--prefill-min-tokens` wide (default 32) run on the device and a shorter tail runs on the CPU, so a device graph and a CPU graph never share a shape. Speculative decoding widens CPU graphs and runs a second context over the same weights, so the two are refused together for now. ### The arena: two layers of the model at a time A model larger than RAM cannot give the NPU a copy of itself. The NPU gets two layer-sized slots instead, and every layer's weights are bound to one of them, alternating. While the NPU computes layer k out of one slot, loader threads fill the other with layer k+1: - **experts** are read from the gguf with `O_DIRECT`, one expert at a time, and handed to the backend with `ggml_backend_tensor_set` on a per-expert view. That call is where the Hexagon backend repacks them into its matrix-engine tiles, so the repack runs in parallel across the loaders; - **the other layer weights** (attention, norms, shared experts) are already resident on the host, so filling them is a copy, not a read. Pacing uses points the graph already offers. The experts wait at the layer's routing node, which the streamer knows how to isolate; there the arena reads the routed ids and loads what is missing (below). The other weights are needed before the routing, so they wait at the last node of the previous layer, which the capture pass learns per layer because no node name is common to every architecture. A graph that skips a pacing point fails the decode rather than compute on a slot that never filled. Measured: the slots cost about 900 MB for the model above, instead of the 21 GB the model is. ### Reading only the routed experts (default; `--no-prefill-routed` reads whole layers) Loading every expert assumes a wide graph routes to nearly all of them. A phone agent's prompt does not: measured with `--decide-probe` on the same model with top-4 routing, a 130 to 480-token prompt routes to about 128 of each layer's 256 experts, and a 55-token one to about 64. A layer is read in two parts. Ahead of its routing, while the NPU computes the layer before, the loaders read the experts the previous graph routed at that layer: consecutive decisions of an agent route much alike. At the layer's routing node the arena reads the routed ids (`ggml_backend_tensor_get`, wherever the tensor lives) and queues what the prediction missed ahead of everything else, then waits for it. The expert matmul reads only routed experts, so what the slot holds for the others never reaches the result: the output is the same bit for bit. A layer that routes to more than `--prefill-routed-full` of its experts (0.85 by default) gets the next layer read whole, as `--no-prefill-routed` does everywhere, so a long prompt, which routes to everything, does not pay for a prediction it cannot use. The first graph of a session has no prediction and reads whole layers. Measured on the phone above, Qwen3.6-35B-A3B Q4_0 streamed with the settings of an on-device Android UI agent (`--overlap --io-threads 4 --dense-weights ahwb -t 6 -c 2048 --n-expert-used 4`, ubatch 2048), 35 decisions: 31 over Android screens (129 to 476 tokens) and 4 short questions. Same session, same build, whole layers and routed: | | whole layers | routed | |---|---:|---:| | prefill, median | 7.68 s | 4.16 s | | arena reads, median | 17.0 GiB | 9.2 GiB | | graph waiting on the arena, median | 6.4 s | 2.9 s | | `choice_logp` identical | | 35 of 35 | The prediction missed 11% of the routed experts. It misses more when the prompt changes domain (the short questions after the screens: about 40%), and those prompts still ran in 2.6 to 2.9 s. Other models on the same phone, same settings, 18 of the screen decisions each, every answer identical between the two: | model | experts routed per layer | whole layers | routed | | |---|---:|---:|---:|---:| | Qwen3.6-35B-A3B Q4_K_M | ~50% of 256 | 9.95 s | 5.37 s | 1.85x | | Gemma 4 26B-A4B Q4_K_M | ~58% of 128 | 6.69 s | 3.70 s | 1.81x | | Nemotron 3.5 30B-A3B Q4_0 | ~79% of 128 | 7.42 s | 6.81 s | 1.09x | The gain is what each layer leaves unrouted: Nemotron's prompts route to most of its experts, many of its layers cross the 0.85 fallback, and it barely gains. The figures are for top-4 routing; with a model's own top-8 a prompt routes to more experts and the gain is smaller (not measured). Per-decision numbers: `bench-data/2026-09-29-prefill-routed/`. ### What else had to move - **The KV cache and recurrent state** move once, at load, into the NPU's host buffer type: memory the CPU reads directly and the NPU addresses too, so decode and prefill share one cache. Left in plain CPU memory, every attention of a device graph would run on the CPU. The Hexagon backend exposes that buffer type only with `GGML_HEXAGON_HOSTBUF=1`, which the app sets. The buffers llama.cpp first allocated the state in stay allocated, since only llama.cpp can free them, but nothing reads them after the move, so their pages are handed back to the kernel (and again after each `llama_memory_clear`, which rewrites them). Kept resident they would double the KV cache: 1760 MiB on Gemma 4 26B-A4B at an 8192-token context, where Qwen3.6, mostly linear attention, has 143 MiB. - **Weights that are not a matmul's matrix.** A backend may hold a `WEIGHTS` buffer in a form only its matmul kernels address: with DMA64 on (the default above Hexagon v79), the Hexagon backend maps such a buffer for DMA only, and most of its other kernels refuse it at run time, which aborts the graph. Gemma 4 met it first: it broadcasts a per-expert scale with `REPEAT`. The capture pass records every layer weight some op reads other than as the matrix of a `MUL_MAT`/`MUL_MAT_ID` (through views too), and those go to a second pair of slots in plain device memory, in their own type. The matrices keep the `WEIGHTS` slots. No op moves: the scheduler still runs each on the device, now on memory every kernel can read. - **Weight types the device refuses.** A "Q4_0" gguf is a mix: the one above keeps its shared experts in Q5_0 and four attention projections in Q6_K, which the Hexagon matmul does not take. Left alone they ran on the CPU inside every device graph (65 matmuls, 111 splits, 17.6 s instead of 11.2 s). Such a weight is carried to the device in the nearest type it takes (Q8_0 for a quantised grid, F16 or F32 for a float one), converted once at load. The file and the CPU decode are untouched. Which type is asked of the device with a probe matmul, not read off a list. - **Compute buffers.** llama.cpp reserves compute memory at load for the widest graph with every weight on the CPU, a graph this session never runs there, and it reserves a logit row per token of the ubatch, each the width of the vocabulary: 2.2 GB at ubatch 2048, which pushed decode into thrashing (0.5 tok/s). The reservation is redone with the weights on the device, and the logit rows are capped at 128, which brought the CPU's buffer to 154 MB. ## Requirements - A model whose expert tensors the NPU's `MUL_MAT_ID` takes: Q4_0, Q4_1, Q8_0, IQ4_NL, MXFP4, and since the 2026-09-26 llama.cpp base the K-quants Q4_K, Q5_K and Q6_K, which is what a Q4_K_M is made of. A dense weight in any other type is converted for the device (see above); Q3_K and below are not taken for experts. gpt-oss is natively MXFP4. - The Hexagon backend in the build: `scripts/build-hexagon-android.sh` builds the CLI, the backend and one DSP-side skel per NPU generation (v73 to v81) inside upstream's Snapdragon toolchain container, and `scripts/stage-hexagon-jnilibs.ps1` stages them into the app. The release APK is built the same way by CI, so it carries all of them. - `--prefill-loaders N` (default 8) sets the threads that fill the slots, apart from `--io-threads`, which stays the decode's read lanes. Each loader reads and repacks, and a K-quant repack is CPU-heavy: on a Q4_K_M, 4 loaders left the NPU waiting 10.3 s of a 15.1 s prefill, 8 left it 5.0 of 10.0. - On device: `ADSP_LIBRARY_PATH` pointing at the directory with the `libggml-htp-v*.so` skels (fastrpc resolves the one for the phone's NPU through it), and `GGML_HEXAGON_HOSTBUF=1`. `GGML_HEXAGON_OPPOLL=1` makes the host poll for the DSP instead of waiting on an interrupt, which halves the cost of each crossing. When the device is not there the run does not fail: a name the registry does not know (no backend in the build, or a phone without the fastrpc driver, where Hexagon registers nothing) and a device that does not open (a Snapdragon older than v73, which registers and then refuses a session) both leave the whole run on the CPU, with a `bmoe:` line on stderr saying which, and `prefill_dev_tokens` stays 0. The device is opened once before the load to find out, and the context reuses that session. Without `--prefill-device` a Hexagon build keeps the NPU out of the run altogether. llama.cpp, given no devices, lists every GPU-type device and opens a backend on each, and Hexagon reports itself as one; the engine drops any such device that can reach neither a host buffer type nor host pointers, since with no layer assigned it could do nothing but open a DSP session. See `docs/seam.md`. ## Correctness Gates G16 and G17 run the whole path against a loopback `rpc-server` fronting the CPU: the same kernels, so placement is the only difference, and output, perplexity and the bytes the arena reads must all match an all-CPU run bit for bit, with and without a cache, with a slowed loader, and across several generates in one session. Removing either wait in the arena fails them (checked). G17f runs the routed arena the same way (fallback disabled, so every layer after the first graph is read from a prediction plus the routing node) and requires fewer bytes read; G17g skips the routing-node reads and requires the perplexity to change, so G17f cannot pass by luck. The RPC backend is only that test fixture: it is built with the tests alone, on 127.0.0.1, and neither the CLI nor the app accepts an RPC endpoint. On the NPU itself the matrix engine computes in fp16, so the output is not bit-identical to the CPU's. Price it with `--ppl` on the same Q4_0 model with and without the flag before relying on it. ## Telemetry `BMOE_DONE` and the CSV trailer carry `prefill_dev_tokens`, `prefill_dev_nodes` (nodes the device actually computed), `prefill_dev_read_mib` and `prefill_dev_stall_s`; `BMOE_DECIDE` carries the same device counters plus, in routed mode, `prefill_dev_routed` and `prefill_dev_demand` (experts routed, and those read at the routing node). See [telemetry.md](telemetry.md). ## Why not decode Measured before this feature: the NPU runs the isolated q4_0 matmul 2.4x faster than the CPU at batch 1 and 20x at 512, and still made the decode token slower, 11% over adb and 31% in the app. With experts streamed on the CPU, a token crosses to the device and back about twice per layer, 91 times on a 40-layer model, at 0.3 to 0.5 ms each. Putting whole layers on the NPU removes the crossings, but on unified memory those layers take RAM from the expert cache one for one, and it lands at parity. A prefill graph crosses the same boundaries once per thousands of tokens instead.