# Inference
## Install
```bash
hf download startlux-models/StartLux-Decision-4B --local-dir StartLux-Decision-4B # or any size, see below
# or, from ModelScope: modelscope download StartLuxAI/StartLux-Decision-4B --local-dir StartLux-Decision-4B
pip install -r requirements.txt
```
The six models are in the [StartLux-Decision collection](https://huggingface.co/collections/startlux-models/startlux-decision-6abba92b301b573fa154d493) on Hugging Face: StartLux-Decision-0.8B, 2B, 4B, 9B, 27B and 35B-A3B, under
`startlux-models/`. The same repositories are on ModelScope under [`StartLuxAI/`](https://modelscope.cn/organization/StartLuxAI). The examples below use a
local folder called `StartLux-Decision-4B`.
`requirements.txt` includes `flash-linear-attention` and `causal-conv1d`. They matter more than anything else on this
page. The models use linear-attention layers, and without these two packages transformers quietly falls back to a
plain PyTorch implementation that is more than ten times slower. Nothing errors; it is just slow. `causal-conv1d` builds
against your CUDA and PyTorch, and if pip ends up compiling it, add `--no-build-isolation`. On Apple Silicon both are
skipped and mlx-lm is installed instead; see [Apple Silicon (MLX)](#apple-silicon-mlx).
Check that the fast path is really on:
```bash
python -m startlux_decision.check StartLux-Decision-4B # fast kernels: active
```
On a machine with a GPU it must say `active`. On CUDA, `StartLuxDecision(...)` refuses to start when the kernels are not
active; set `STARTLUX_ALLOW_SLOW=1` if you really want to run without them. On a CPU-only machine the check is skipped and
everything runs, slowly.
## Serving
```bash
python -m startlux_decision.server --model StartLux-Decision-4B --port 8090
curl -s localhost:8090/health # {"status": "ok", "model": "StartLux-Decision-4B", "fast_kernels": true}
```
The server speaks the TypeSafe `/v1/systemone` format and serves one request at a time per GPU. Run one server per GPU
and put a load balancer in front if you need more throughput. The official TypeSafe SDK works against it:
```bash
export TYPESAFE_BASE_URL=http://127.0.0.1:8090
export TYPESAFE_API_KEY=unused
```
## Request and response
A request has a `state` (a string or any JSON value) and `questions`, each with a `type`:
| type | criteria | answer |
|:---:|:---:|:---:|
| `choice` | `{option: description or null}` | `choice`, `confidence`, `probabilities` over the options |
| `noul` (yes/no) | optional `{"true": ..., "false": ...}` | `noul` = probability of yes |
| `score` | a list of levels, lowest first | `score` (probability-weighted mean level), `confidence`, `legend`, `probabilities` keyed `"0".."n-1"` |
For the support-ticket request in the README, StartLux-Decision-4B answers (numbers rounded here):
```json
{"answers": {
"team": {"type": "choice", "choice": "billing", "confidence": 0.816,
"probabilities": {"billing": 0.877, "shipping": 0.005, "technical": 0.118}},
"urgent": {"type": "noul", "noul": 0.635},
"severity": {"type": "score", "score": 1.335, "confidence": 0.002,
"legend": {"0": "cosmetic", "1": "annoying", "2": "blocks the customer"},
"probabilities": {"0": 0.116, "1": 0.433, "2": 0.451}}},
"usage": {"input_tokens": 290, "output_tokens": 0}, "model": "StartLux-Decision-4B", "latency_ms": 15.7}
```
Each question is rendered as its own prompt with the full state, options lettered A, B, C and so on in the order you
give them, and the answer is read from the logits of those letters. A choice question can have up to 26 options in one
pass. Longer lists are split into near-equal groups of up to 25, the top three of each group go to a final round, and
the options that miss the final keep a small share of probability in proportion to their group score, so every option
still gets a non-zero probability.
Temperatures are per question type and live in `decision_config.json`. Changing them never changes which option wins.
`confidence` follows TypeSafe's definitions, so code that thresholds it behaves the same against either service. For a
choice it is (p_max − 1/n) / (1 − 1/n): 0 for an even split over the n options, 1 when one option has all the
probability. For a score it is 1 minus the probability-weighted distance, in levels, from the most likely level,
divided by the same distance for an even spread measured from the middle level, and floored at 0: probability on a
neighbouring level lowers it less than probability at the far end. Yes/no answers have no `confidence`; |2p − 1| is the
equivalent. Until 2026-10-03 the package returned the top probability as `confidence`; it is still in `probabilities`.
Every response carries an `x-typesafe-request-id` header, `GET /v1/models` lists the model with its release date, and
an error comes back as `{"error": message, "detail": [{"loc", "msg", "type"}]}`, with status 400 for a malformed body
and 422 for a request the model cannot answer, such as a score with more than ten levels.
## Python
```python
from startlux_decision import StartLuxDecision
m = StartLuxDecision("StartLux-Decision-4B") # device defaults to cuda when available
answers, usage = m.decide(state, questions)
answers_list = m.decide_batch([(state1, questions1), (state2, questions2), ...])
```
`decide` is the latency path. `decide_batch` is the throughput path: every question of every request is sorted by
length and packed into padded forward passes of up to `max_batch_tokens` tokens (65,536 by default). On a random 2%
sample of the Decision Index suite (2,678 requests) StartLux-Decision-4B took 140 s with `decide_batch` against 337 s calling
`decide` once per request, and the chosen options agreed on 99.94% of the 6,897 questions. The differences are bf16
rounding between batch shapes; on 981 JevBench and Typed Decisions questions the largest probability difference was
0.008.
## Long inputs
A prompt can be as long as 262,144 tokens, the native context of all six models (`max_length`, or `--max-length` on the
server). Every question of a request carries the full state, so their prompts begin with the same tokens. When that
shared beginning is longer than 4,096 tokens it runs once: the model reads it in chunks of 32,768 tokens
(`prefill_chunk`) and keeps what each layer carries forward, the keys and values of the attention layers and the
convolution and recurrent state of the linear-attention layers. Then all questions run as one batch on top of it. A
single question over a long document runs the same way, so memory stays close to the keys and values of the prompt
instead of growing with the activations of the whole input. Attention over the stored keys runs as fused cuDNN or flash
attention calls whose outputs are merged by their log-sum-exp, without an attention mask. Prompts that long take seconds
rather than milliseconds, and at 256K tokens the 27B and the 35B-A3B need most of an 80 GB GPU.
## Images
The checkpoints include a vision tower, and the package uses it. Images are part of the evidence of a request:
```python
answers, usage = m.decide("Photo taken at delivery: ",
{"damaged": {"type": "noul", "instructions": "Is the parcel damaged?"}},
images=["parcel.jpg"])
```
```bash
curl -s localhost:8090/v1/systemone -H 'Content-Type: application/json' -d '{
"state": "Photo taken at delivery: ",
"questions": {"damaged": {"type": "noul", "instructions": "Is the parcel damaged?"}},
"images": ["data:image/jpeg;base64,'"$(base64 < parcel.jpg | tr -d '\n')"'"]
}'
```
- In Python an image can be a PIL image, a file path, encoded bytes, a base64 string or a data URI. Over HTTP, `images`
is a list of base64 strings or data URIs.
- `` in a string state marks where each image goes, one mark per image, in order. Without marks the images come
first (as `Image 1:`, `Image 2:` ... when there are several), followed by the state.
- Each image is resized to at most 1,048,576 pixels, keeping its aspect ratio (`max_pixels`, `--max-pixels`), and
becomes one token per 32 × 32 pixels: a 1024 × 1024 photo is 1,024 tokens.
- The vision tower runs once per request, and every question of the request reads the same image tokens. Image
requests run on the eager path, without the CUDA graphs.
- The vision tower adds 0.2 to 0.9 GB of GPU memory. `StartLuxDecision(..., images=False)` or `--no-images` leaves it
out.
- The MLX and GGUF backends read text only.
## Why it is fast
Three things, in order of how much they matter.
1. The fast kernels above. Without them everything else is moot.
2. No generation. One forward pass per request, each question one row of the batch, and only 26 rows of the output
matrix are ever multiplied.
3. CUDA graphs. At start-up the model records graphs in a single shared memory pool: one per padded input length
(128, 192, 256 ... 4096 tokens) for a single question, and one per question count and length for requests with two
to four questions of up to 1024 tokens (40 graphs in all). The experts of StartLux-Decision-35B-A3B run as grouped
matrix multiplications, with no synchronisation with the host, so they are recorded as well. A short request is right-padded to the next length and
all its questions are replayed as one graph, which removes the per-layer kernel launch overhead that dominates small
inputs. Padding goes after the last prompt token, every layer is causal and the rows never mix, so it never affects
the position that is read. Recording adds to start-up time; set `STARTLUX_GRAPHS=0` to skip it. Requests with more
or longer questions run on the eager path, still as one batch, and bulk work belongs in `decide_batch`.
Eager and batched inputs are padded to a fixed ladder of lengths, because the linear-attention kernels are compiled
once per sequence length. At start-up the server runs one request through both the graph and the eager path and
prints the largest probability difference (0.0 for all six models); above 0.02 it drops the graphs.

Measured end to end over HTTP on one H200, bf16, one request at a time (`eval/latency.py`, 20 warm-up and 200 timed
requests). "3 fields" is one choice, one yes/no and one score question on a support ticket; the three prompts total
891 tokens because each question carries the full state, and the three run together in one forward pass, as
Intern-Decision's fields do.
| Model | 3 fields: mean | P50 | P95 | one yes/no question |
|:---:|:---:|:---:|:---:|:---:|
| StartLux-Decision-0.8B | 12.2 ms | 12.2 ms | 14.2 ms | 8.3 ms |
| StartLux-Decision-2B | 15.5 ms | 15.5 ms | 17.2 ms | 9.6 ms |
| StartLux-Decision-4B | 26.0 ms | 26.0 ms | 27.3 ms | 14.7 ms |
| StartLux-Decision-9B | 35.7 ms | 36.2 ms | 37.4 ms | 17.6 ms |
| StartLux-Decision-27B | 102.3 ms | 102.5 ms | 104.3 ms | 50.7 ms |
| StartLux-Decision-35B-A3B | 52.5 ms | 52.9 ms | 54.5 ms | 36.4 ms |
| StartLux-Decision-4B, graphs off | 90.3 ms | 89.5 ms | 92.5 ms | 87.5 ms |
For reference, Intern-Decision reports 34.0 ms (0.8B), 33.3 ms (2B) and 44.2 ms (4B) on an RTX 4090 for a request of the
same shape. Different hardware, so read it as a ballpark, not a head-to-head. For Jev 1.13 we sent the same two requests
to the TypeSafe API 100 times each on 2026-09-29. Its gateway reports 64.0 ms of server time for the three fields and
63.5 ms for the single yes/no question (`x-envoy-upstream-service-time`, means; medians 57.5 and 58.0 ms); Jev answers all
fields at once, so the count barely matters. End to end from our cluster the requests took about 330 ms, of which about
260 ms is the network. Intern-Decision's 109.7 ms for the Jev API is also an end-to-end number.
## FP8
The models also run with FP8 weights and activations (torchao, `Float8DynamicActivationFloat8WeightConfig` with per-row
scales on the text backbone; the letter readout stays in bf16). On 2,541 decisions (JevBench public, Typed Decisions
test, ToolACE test) the chosen option matched bf16 in this share of cases:
| Model | agreement with bf16 | accuracy bf16 -> FP8 |
|:---:|:---:|:---:|
| StartLux-Decision-0.8B | 96.3% | 78.00 -> 77.76 |
| StartLux-Decision-2B | 96.5% | 80.44 -> 80.05 |
| StartLux-Decision-4B | 98.0% | 82.37 -> 81.98 |
| StartLux-Decision-9B | 98.2% | 83.67 -> 83.23 |
| StartLux-Decision-27B | 98.6% | 82.96 -> 82.76 |
FP8 is a memory option here, not a speed option. Weight memory roughly halves (StartLux-Decision-27B takes 27.5 GiB on the GPU
after quantisation), but in this test FP8 was about 2.7 times slower per decision than bf16, both run eagerly with one
question per forward pass: for inputs this short, quantising activations on the fly costs more than the smaller
matmuls save.
One caveat: with dynamic per-row activation scales, an all-zero padding row gets a zero scale and produces NaN, which
then leaks into real rows through attention. When we let FP8 run on padded batches, agreement with bf16 fell to about
52%. Run FP8 one question per forward pass, without padding.
## GGUF and llama.cpp
GGUF files of every size are on Hugging Face, one repository per size and precision (BF16, Q8_0 and Q4_K_M), named
`startlux-models/StartLux-Decision---GGUF` (on ModelScope: `StartLuxAI/` with the same names). They hold
the text decoder only. The prompt format, the
option-letter readout and the per-type temperatures stay in this package, and `startlux_decision.gguf_server` puts them
in front of llama-server:
```bash
hf download startlux-models/StartLux-Decision-4B-Q8_0-GGUF --local-dir StartLux-Decision-4B-Q8_0-GGUF
cd StartLux-Decision-4B-Q8_0-GGUF
pip install -r requirements.txt # transformers and torch; a CPU build of torch is enough
llama-server -m StartLux-Decision-4B-Q8_0.gguf -ngl 99 -c 16384 --parallel 4 --port 8081
python -m startlux_decision.gguf_server --model-dir . --llama http://127.0.0.1:8081 --port 8090
```
The server speaks the same `/v1/systemone` format as `startlux_decision.server`. It asks llama-server for the
next-token log-probabilities at the answer position and applies the same readout, temperatures and wide-choice rounds,
so a GGUF file can be compared with the original weights question by question; the GGUF table in the README does that
on the public JevBench items. Plain chat with a GGUF file does not give these decisions. llama.cpp has to be recent
enough to support this model (build b10454 or newer).
## Apple Silicon (MLX)
On a Mac, `pip install -r requirements.txt` installs [mlx-lm](https://github.com/ml-explore/mlx-lm) instead of the CUDA
kernels, and the same commands run the model with MLX (`--backend auto`, the default; `--backend torch` forces PyTorch):
```bash
hf download startlux-models/StartLux-Decision-4B --local-dir StartLux-Decision-4B
pip install -r requirements.txt
python -m startlux_decision.server --model StartLux-Decision-4B --port 8090 # bf16
python -m startlux_decision.server --model StartLux-Decision-4B --port 8090 --int8 # M5 and later
```
```python
from startlux_decision.mlx_model import MLXDecision
model = MLXDecision("StartLux-Decision-4B") # int8=True on M5 and later
answers, usage = model.decide(state, questions)
```
The MLX backend does three things:
- **The forward pass runs in MLX.** `MLXDecision` replaces only the forward pass, as the GGUF server does, and uses
mlx-lm's implementation of the model with its Metal kernels for the linear-attention layers. PyTorch on MPS has no such
kernels and runs reference code. Prompts, the option-letter readout, the temperatures and wide choices are the
package's own code. In bf16 the public JevBench scores of 0.8B, 2B, 4B and 9B are the published ones.
- **Prefixes are reused.** Every prompt starts with the same system text. It runs once when the server starts, and each
request continues from its cache (attention keys and values, linear-attention conv and recurrent states). The
questions of a request are rows of one forward pass. When they share a long piece of evidence, the evidence runs once,
and a later request about the same evidence runs only its questions.
- **`--int8` uses int8 matmuls on M5 and later.** The large projections run as int8 × int8 matmuls on the GPU's neural
accelerators. Activations are quantized per token and weights per channel. SmoothQuant scales, computed at start-up
from a few built-in requests, are applied first. On the public JevBench items the decisions match bf16 (the 8-bit
model for 27B) on 97.4 to 99.6% of the items.
StartLux-Decision-27B in bf16 needs 56 GB for its weights. On a 64 GB Mac, convert it to 8-bit first. The 8-bit model
scores 208 of 231 on public JevBench; bf16 scores 209.
```bash
mlx_lm.convert --hf-path StartLux-Decision-27B --mlx-path StartLux-Decision-27B-MLX-8bit -q --q-bits 8
cp StartLux-Decision-27B/decision_config.json StartLux-Decision-27B-MLX-8bit/
python -m startlux_decision.server --model StartLux-Decision-27B-MLX-8bit --port 8090 --int8
```