--- name: processor-model-parity description: Add or verify a model's chat prompt rendering and tokenization in rust/sglang-processor so it matches SGLang's Python serving path exactly (same prompt text, same token ids). Use when adding a model to sglang-processor, porting a model's rendering from sgl-router, bumping Dynamo's renderer, or debugging a processor-vs-Python prompt mismatch. --- # Processor model parity `rust/sglang-processor` renders chat requests for SGLang's Rust hosts (the Rust server, sgl-router, the renderer). A model is supported only when, for the same request body, the processor produces **the same prompt text and the same prompt token ids** as Python's `OpenAIServingChat`. Python is the reference, Dynamo is the implementation, and SGLang code exists only where the two differ. Scope: the request -> prompt -> token ids path. Output parsing (reasoning and tool calls) and multimodal preprocessing are not covered here. The worked example is DeepSeek-V4-Flash-0731: `src/render/models/deepseek_v4.rs` plus `tests/fixtures/parity/deepseek-v4-flash-0731.json`. Paths below are relative to `rust/sglang-processor/` unless they start with `python/`. ## Where a model's code goes | Piece | File | Needed when | |---|---|---| | Identity -> formatter | `src/render/selection.rs` (`select_chat_formatter`) | always; mirror the order of `chat_encoding.resolve_chat_encoding_spec` | | Formatter variant | `src/render/mod.rs` (`ChatFormatter`, its `render_request` arm, plus the `render_prompt`, `stop_strs` and `resolve_thinking` arms) | always | | SGLang-specific steps | `src/render/models/.rs`, re-exported in `models/mod.rs` | only where Dynamo diverges from Python | | Parity cases | `tests/fixtures/parity/.json` | always | | README row | the `render` table in `README.md` | always | No new test code: `tests/parity.rs` checks every fixture in the directory. `render_request(Value, ..)` takes the SGLang request **body**, not Dynamo's `OAIChatLikeRequest`, because Python reads fields the trait lacks (`task`, `continue_final_message`, the `reasoning` object). Pick the shape by how Python renders the model: - **Python calls its own encoder** (a non-`None` spec such as `dsv4` or `dsv32`): give the model a dedicated `ChatFormatter` variant and `render_request` arm, like DeepSeek-V4. If `selection.rs` already wires the model to a Dynamo `PromptFormatter`, replace that branch. - **Python calls `apply_chat_template`** (Jinja, spec `None`): the first such model adds one generic arm that adapts the body into `OAIChatLikeRequest`, in its own PR. sgl-router's `ChatRequest` in `experimental/sgl-router/src/tokenizer/chat_formatter.rs` is a working adapter. ## Method 0. **Pick the checkpoint.** Use the one the model's cookbook page recommends (`docs/cookbook/`), pin its Hugging Face commit as `revision`, and cache it. Note whether its tokenizer ships a `chat_template`; that can change the spec. 1. **Trace the Python path for the model.** Write down every transformation between the HTTP body and `prompt_ids`, in order: - `protocol.py::ChatCompletionRequest.normalize_reasoning_inputs`: `reasoning` and `reasoning_effort` set the default `thinking` / `enable_thinking`. Reuse `render/reasoning.rs` for this rather than porting it again. - `serving_chat._convert_to_internal_request`: `chat_template_kwargs.reasoning_effort` replaces the request effort; server `--default-chat-template-kwargs` merge in. - `chat_encoding.resolve_chat_encoding_spec`: which branch renders the model (a spec such as `dsv4`, `dsv32`, `kimi_k3`, `inkling`, or `None` for Jinja). It also reads the server's `--tool-call-parser` and whether the tokenizer has a chat template. If the deployment depends on the parser, set `tool_call_parser` in the fixture, and add it to `ChatFormatterOptions` in the same PR. Note any per-checkpoint state the server resolves once (DeepSeek-V4's effort profile reads `encoding/encoding_dsv4.py`). - That spec's branch in `serving_chat._apply_jinja_template` / `_encode_messages`: message dumping, content flattening, `_handle_last_assistant_message`, system insertion, tools, model-specific fields, the encoder or `apply_chat_template` call, and how the prompt is tokenized. - The encoder itself (for example `encoding_dsv4.py`), line by line. History rules live there: dropped turns, `` on empty reasoning, argument formatting, and the errors it raises. This list is the **divergence inventory**. 2. **Find Dynamo's counterpart.** Check `dynamo_renderer::native_formatter_for`, the model's low-level encoder (for example `deepseek::v4::encode_messages_with_options`), or the Jinja template. Prefer the lowest-level Dynamo entry that lets SGLang own request semantics. Dynamo's OpenAI-level formatters apply their own defaults (they filter tools by `tool_choice`, inject `response_format`, and map effort), which often differ from SGLang's. Diff the low-level encoder against Python's for the rules in step 1. 3. **Write one request case per inventory item** in the fixture (`name` + `request` only), then generate the expected outputs from Python (below). Cover the checklist; reuse the request shapes from the DeepSeek-V4 fixture where they apply. 4. **Implement.** Start from a plain Dynamo call and run `tests/parity.rs`. For each failing case, add the smallest SGLang step that fixes it, with a one-line doc comment naming the Python source it mirrors. Never patch the fixture to match Rust. 5. **Verify** text offline, then token ids with the checkpoint cached (commands below). If sgl-router or the renderer uses the formatter, run their suites too. ## Checklist (DeepSeek-V4 case that covers each) - **Message dump**: roles lowercased; unknown and null fields dropped; `user` reduced to role and content; null content becomes `""`. (`ignored_fields`, `tool_results`) - **Content parts**: which part types survive and the join separator. (`parts`) - **Tools**: all request tools or only those `tool_choice` allows; field order and defaults of `Function.model_dump`; message-level tools; empty lists. (`tools`, `tools_none`, `tools_named`, `message_tools`, `empty_message_tools`) - **Pydantic coercion**: typed request fields are coerced before rendering. For example, `strict: 1` and `defer_loading: "false"` become booleans, and `strict: null` is rejected. Probe the pydantic model in Python to get its exact rules; don't guess them. (`tool_bool_coercion`, `tool_strict_null`, `continuation_coerced`) - **Tool-call arguments**: how the encoder wants them. DeepSeek-V4 takes a compact JSON string; other encoders format them their own way. Keep key order (`serde_json` `preserve_order` is declared in `Cargo.toml`). Floats must print as Python's `json.dumps` prints them (`1e-06`, `1e+16`, `100.0`). (`tool_results`, `agentic_thinking`, `tool_float_values`, `history_float_arguments`) - **Thinking**: kwargs `thinking` > request effort (`!= "none"`) > `reasoning.enabled` > server default kwargs > `SGLANG_DEFAULT_THINKING`. `enabled` follows Python truthiness (`1` is on); strings use Python's yes-word set. (`thinking_false`, `thinking_none`, `drop_thinking_ignored`, `reasoning_enabled_numeric`, `reasoning_enable_string`) - **Effort**: precedence between kwargs, `reasoning`, `reasoning_effort` and env; the checkpoint's profile mapping. (`effort_high`, `effort_max`, `kwargs_effort`, `effort_conflict`, `reasoning_object`) - **Final assistant turn**: becomes a user turn, or with `continue_final_message` a separately tokenized prefix. Check whether content is flattened before the split (it is for DeepSeek-V4). Run each check where Python runs it: a final assistant turn's tool calls are discarded, so object checks on their arguments belong after the split. Hosts reach the model through `render_prompt` too, so check the renderer's continuation handling (`sglang-renderer` `render()`) as well. (`final_assistant`, `continuation_parts`, `continuation_bos`, `continuation_only`, `final_tool_call_no_arguments`, `continuation_tool_call_null_arguments`) - **Model-specific fields and turns**: `task` placement, a system turn mid-conversation, inserted empty system turns. (`task_action`, `task_after_developer`, `consecutive_task`, `mid_system`) - **History**: the encoder's rules for earlier turns: reasoning kept or dropped, whole turns dropped, closing tags on empty reasoning. (`multi_turn`, `thinking_multi_turn`) - **Tokens**: Python's `tokenizer.encode` adds special tokens by default; the continuation prefix is encoded alone and loses a leading BOS (`_append_assistant_prefix_to_prompt_ids`). `tests/parity.rs` does both, using the fixture's `bos_token_id`; `render_prompt` returns the prefix as its own segment so hosts keep that boundary. (`continuation_bos`, `continuation_after_text`) - **Errors**: reject what Python rejects. A case Python raises on is recorded with `error`, and `tests/parity.rs` then requires `render_request` to fail. That includes pydantic rejections and Python crashes (a 500 is still a rejection). (`invalid_tool_arguments`, `continuation_only`) If Dynamo rejects a shape Python accepts, say in the PR that hosts fall back to the engine for it. - **Known gaps**: when a mismatch is inside Dynamo and out of the processor's reach, fix it upstream and mark the case `known_gap` with the reason. The test reports the case without failing, and fails once it matches so the marker gets removed. (`tool_float_values`: Dynamo's `deepseek::common::to_json`) ## Generating fixtures `tests/scripts/generate_parity.py` loads the pinned snapshot and resolves the spec and per-checkpoint state with SGLang's own resolvers. It then renders every case through the real `OpenAIServingChat._apply_jinja_template` and writes: - `prompt`, `token_count` and `token_sha256` for each case, or `error` when Python raises (`name`, `request` and `known_gap` are kept as written). An `AttributeError` aborts instead: the stub server lacks something, so extend it; - the fixture-level `config` (the `config.json` fields the processor reads, plus resolved overrides such as the effort profile) and `bos_token_id`. ```sh cd rust/sglang-processor PYTHONPATH=$REPO/python HF_HUB_CACHE= HF_HUB_OFFLINE=1 \ python tests/scripts/generate_parity.py tests/fixtures/parity/.json ``` When a new spec needs more, extend `serving_chat()` in the script (one place): - **More server state:** set the extra attribute that `__init__` resolves, for example `_dsv41_default_reasoning_effort`. - **Tokenizes through `apply_chat_template(tokenize=True)`:** the recorder sees no text. Record the template's text output instead. Pin `revision` to a commit hash and regenerate only on purpose. ## Verify ```sh cd rust cargo test -p sglang-processor --locked # text, offline HF_HUB_CACHE= cargo test -p sglang-processor --test parity --locked # + token ids cargo clippy -p sglang-processor --all-targets --locked -- -D warnings cargo check -p sglang-processor --no-default-features --features render,tokenizer --locked ``` The token check loads the pinned commit's snapshot from the cache. When the snapshot is missing, the test prints `token ids not checked` (visible with `-- --nocapture`), so confirm that line is absent before claiming token parity. With the snapshot cached, the test also checks that the checkpoint resolves the fixture's recorded DeepSeek-V4 profile. ## Pitfalls - Dynamo version bumps change rendering silently: rerun every fixture after bumping `dynamo-renderer` or `dynamo-tokenizers`. - Env vars (`SGLANG_DEFAULT_THINKING`, `SGLANG_DSV4_REASONING_EFFORT`) are read per request, as in Python. The generator pins them and `tests/parity.rs` clears them, so fixtures do not depend on the machine. - Typed hosts (the renderer's `OAIChatLikeRequest` path) cannot carry every field, such as `task` and message-level `tools`. Parity is defined on `render_request`; report host-adapter gaps separately rather than bending the model code. - Keep model code in its own file: no model-specific branches in `render/mod.rs` beyond the dispatch arm.