# swanOne bench-request submission package Everything needed to run the filing, which is JevBench issue #23. **On a DGX Spark, start with [`SPARK.md`](SPARK.md).** The filing's serve command was written for the benchmark's x86 H100 and will not start on a Spark as written. **On an H100 NVL, [`H100.md`](H100.md)** has the whole route in one place, with the context cap raised to 16384. ## Built on MiaAI Lab's recipe — please read None of this would run without **[MiaAI Lab](https://x.com/MiaAI_lab)**'s [`Qwen3.8-Flash-Next-Single-DGX-Spark`](https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark). Serving a 99 GiB NVFP4 model from one Spark's unified memory is their work: the PLE CPU-offload machinery and the GB10 stream-memory diagnosis behind it, the memory-mapped packed PLE table, the MTP draft-vocabulary index, the MXFP8 kernel fallbacks and the FP8-KV cache path. **Thank you, Mia and team** — the measurements, the failure counts and the willingness to publish the cells that did not work are what made this possible from the outside. To be exact: **the nine files in `patches/` are their patch output, not our engineering.** Their generators reproduce all nine of our files byte-for-byte. What is ours is the packaging and the shim. The shim's approach — reading the answer from the model's probabilities over the option letters in one forward pass — is [NInfer](https://github.com/igorls/ninfer)'s; the code is ours. Their recipe licenses its **generators** under AGPL-3.0-or-later, but it also states that the files they generate "keep vLLM's own Apache-2.0 headers and remain Apache-2.0 works" — and our nine files are that output. So `patches/` is **Apache-2.0**, and **none of their AGPL scripts are redistributed here**. `shim/` is ours and is MIT. That reading is ours, from their README — we have not agreed it with them directly. `LICENSE-NOTICE.md` has the detail and `CREDITS.md` names everyone whose work is here. patches/ the nine patched vLLM files, each mounted over an absolute path in the image, with the pre-patch originals so the diff is reproducible patches/MANIFEST.md targets, changed-line counts, sha256s, apply command patches/*.orig pre-patch originals for the 5 top-level files patches/orig/ pre-patch originals for the 4 ple_offload files assets/ draft_vocab_en_code_47k.txt — 47,172-token MTP draft vocabulary (optional; see §4.6) shim/ typesafe_native_shim.py — implements /v1/systemone over vLLM logprobs shim/test_shim.py its answers and status codes, checked without a GPU rescore/ the filed run, per item, and its v1.3 rescore (see rescore/README.md) baselines/ the same model writing its answer out, reasoning on and off (see baselines/README.md) H100.md the whole route on an H100 NVL · SPARK.md the route on a DGX Spark Order of operations: patches -> server (filing §4.2, or H100.md) -> shim (filing §3) -> harness (filing §5). Nothing here needs credentials. Do not commit a HuggingFace token into this tree. ## What one decision costs Same model, JevBench's public items, tokens generated per decision: | how the decision is answered | tokens generated per decision | |---|---:| | reasoning on, answer written out | 714.3 on average — 13 of 173 stopped at the 4,096-token cap | | reasoning off, answer written as JSON | 49.9 on average | | the one-token readout in `shim/` | 1 — 231 of 231 | The prompt (about 700 tokens) is read in every case; the readout removes the generation. Re-derive all three with `python3 baselines/tokens.py`. ## Where it landed, category by category JevBench publishes per-item outcomes for the systems it has run on the same 231 public items. The filed run is lined up against 40 of them, item for item: | slice | items | ours | systems ahead | tied | max of the 40 | |---|---:|---:|---:|---:|---:| | all public items | 231 | 204 | 2 | 0 | 226 | | easy | 48 | 48 | 0 | 28 | 48 | | standard | 72 | 70 | 6 | 4 | 71 | | hard | 111 | 86 | 2 | 0 | 107 | | category | items | ours | systems ahead | tied | max of the 40 | |---|---:|---:|---:|---:|---:| | adequacy | 12 | 10 | 12 | 3 | 12 | | adversarial | 6 | 6 | 0 | 14 | 6 | | ambiguous | 7 | 6 | 5 | 3 | 7 | | extraction | 24 | 24 | 0 | 23 | 24 | | fact | 12 | 12 | 0 | 29 | 12 | | intent | 24 | 24 | 0 | 16 | 24 | | judge_hard | 17 | 12 | 14 | 1 | 16 | | long_policy | 19 | 16 | 2 | 0 | 19 | | multi_hop | 18 | 15 | 3 | 3 | 18 | | ordinal | 12 | 12 | 0 | 24 | 12 | | policy | 12 | 12 | 0 | 8 | 12 | | probability | 10 | 8 | 4 | 2 | 10 | | routing | 12 | 12 | 0 | 16 | 12 | | routing_hard | 5 | 5 | 0 | 23 | 5 | | temporal_numeric | 15 | 6 | 6 | 5 | 14 | | tool_selection | 12 | 12 | 0 | 37 | 12 | | tradeoff | 6 | 5 | 4 | 1 | 6 | | trap | 8 | 7 | 16 | 6 | 8 | The two systems ahead overall and on the hard tier are DeepSeek V4.1 Flash (226 of 231) and GPT-5.6 Luna at low reasoning effort (225), both frontier API models, and both at 107 of 111 on the hard tier. This is accuracy on the public items, not the board's Intelligence axis, which weights the tiers, corrects for chance and includes the judge tier. We developed the readout on these same items, so our side of every row is an upper bound; the 40 are JevBench's own runs. One item moves a small category by 4–20 points. Re-derive the tables with `python3 rescore/category_split.py /results/v1.2/jevbench-v1.2-per-task.json`, from the board repository at `fd51755`. ## How to use this repository Everything lives under three directories, and the filing refers to them as placeholders you must substitute before pasting any command. **A literal paste fails**, because `<` is a shell redirect. These placeholders belong to the filing's H100 command; the Spark route in [`SPARK.md`](SPARK.md) needs none of them. | placeholder | set it to | |---|---| | `` | `.../swanone-recipe/patches` | | `` | `.../swanone-recipe/assets` | | `` | a Hugging Face cache directory containing `Mia-AiLab/Qwen3.8-Flash-Next-NVFP4` | | `` | a directory for the PLE table cache (may be empty) | For example, with this repository cloned to `/srv/swanone-recipe`: export PATCHDIR=/srv/swanone-recipe/patches export ASSETDIR=/srv/swanone-recipe/assets `patches/MANIFEST.md` lists, for each of the nine files, its mount target in the image, its changed-line count, and its sha256 — and gives the apply command. The nine files are mounted **over** the image's own copies; they are not baked in. Order of operations: **patches -> server -> shim -> harness.** The server alone does not speak the benchmark's wire format; `shim/typesafe_native_shim.py` is what serves `/v1/systemone`, and it listens on **port 8009**. ## Step 4, written out in full — the filing's §5 was only a flag fragment Filing §5 gives the harness flags but no invocation, which is our error. It is `jevbench`'s own CLI and the dataset is yours to choose, so the complete command is: python3 -m jevbench.cli run \ --tasks /datasets/public/easy.jsonl,/datasets/public/original.jsonl,/datasets/public/hard.jsonl \ --adapter typesafe --endpoint http://127.0.0.1:8009 --key-env '' \ --model swanone --cost-basis self_hosted_gpu --reserve-usd 0 \ --results /results.jsonl \ --raw-dir /raw \ --ledger /ledger.jsonl \ --manifest /manifest.json \ --run-label swanone --delay-s 0 `--endpoint` must match whatever `SHIM_PORT` the shim was started with (§3 uses 8009). Add `--max-model-len`-style limits on your side as your harness requires; nothing in the shim depends on it. **Inputs the shim cannot take.** A prompt over the server's context limit, more than 26 options, or a question type it does not know gets **HTTP 422**, which JevBench's runner scores as one wrong answer and moves past. vLLM's own 401, 403 and 429 pass through unchanged, and any other vLLM failure is a 502; those count toward the runner's rule that three consecutive failures stop the run. Before answering, the shim reads the served model's context limit from vLLM's `/v1/models`; if the server does not list `SHIM_MODEL`, reports no limit, or was started with less than 4096 (`SHIM_MIN_CONTEXT`), it answers 503, so the run stops rather than scoring every longer item wrong. `python3 shim/test_shim.py` checks all of this without a GPU. **Two practical notes from running this ourselves on an H100 NVL:** - **`git` is not installed in the published image; `patch` is.** So the §8.3 route (`patch -p1 -d / < swanOne-vllm-patch.diff`) works inside the container, while a `git clone` does not. - **Run `--max-model-len 16384` (or 8192), not the 4096 in §4.2.** The benchmark describes its hard tier as *"long multi-condition policy documents (2–6k tokens)"*, and *"an input over a system's documented context limit"* counts **wrong**. On an H100 NVL (94 GB) at `--gpu-memory-utilization 0.90`, both caps start with the rest of §4.2 unchanged, and each answered a long-policy test prompt of about 6,360 tokens with the expected label in one token. A larger cap costs concurrency, not memory, and the benchmark sends one request at a time. [`H100.md`](H100.md) has the command with 16384. Nothing here requires credentials. If you would rather have a tarball or a `git diff`, open an issue on the benchmark repository and ask — we will put it wherever is easiest for you.