# MiniMax-H3 Prompt Rewriter for ComfyUI ComfyUI nodes for the [LightX2V MiniMax-H3 T2VA Prompt Rewriter LoRA](https://huggingface.co/lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA). A short prompt goes in; a structured, production-ready audio-video description for [MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3) comes out — entirely locally. [Русская версия](README_RU.md) · [中文版](README_ZH.md) · [Changelog](CHANGELOG.md)

Latest release License Python 3.10+ ComfyUI Registry Hugging Face GGUF adapter, 27B GGUF adapter, 8B GGUF adapter, Omni Video review, in English Video review, in Russian

![The rewriter node in ComfyUI: a short prompt on the left, the structured shot-by-shot description, soundscape and music fields on the right](docs/node_preview.png) The prompt may be in any language the base model reads; the rewrite comes back in English, which is what MiniMax-H3 expects. ```text "A red fox walks through a snowy forest at dawn." + 16:9 + 15s │ ▼ Qwen3.6-27B + Prompt Rewriter LoRA │ ▼ integrated_multimodal_description: [Shot 1] ... [Shot 2] 0:06 ... overall_soundscape: ... non_diegetic_music: ... │ ▼ MiniMax-H3 video + synchronized audio ``` There are three ways to get that output, and the pack ships all of them: | | Rewriter node | Rewriter 8B | Rewriter Omni | Writer nodes | |---|---|---|---|---| | Where the format comes from | the LoRA — a 27B trained until H3 output came out of it unprompted | a second LoRA, on a model that also sees | a third LoRA, on a model that also hears | MiniMax's own writing guide, in the system prompt | | Model | Qwen3.6-27B only | Qwen3-VL-8B-Instruct only | Qwen2.5-Omni-7B only | any instruction-following GGUF | | Smallest working setup | ~10 GB download, ~13 GB VRAM | ~6.1 GB download, ~9 GB VRAM | ~6.2 GB download, ~9 GB VRAM | **2.6 GB download, ~5 GB VRAM** | | Tasks | T2VA | T2VA, I2VA, FL2VA, L2VA | T2VA, I2VA, FL2VA, L2VA, **Ref2VA** | T2VA, I2VA, FL2VA, L2VA, Ref2VA | | Reference frames | described to it in words | **it looks at them** | **it looks at them** | described to it in words | | Clips and sound | described to it in words | described to it in words | **it watches and listens** | described to it in words | | Quality | the reference | the same trained contract, at a third of the download; wobblier on the alignment line | the only one that hears; six fields on Ref2VA | close, and it runs on hardware the LoRA cannot touch | The first three columns are also available as one node — [Universal Rewriter](#minimax-h3-universal-rewriter) — where a tab swaps the adapter and everything else stays where it is. Two of the four read text only. [Reference Caption](#minimax-h3-reference-caption) turns an image, an audio clip or a video into the text they need — 3 to 5 seconds per asset on a 3.4 GB model. When a whole shot's worth of references is waiting, [Multi Reference Caption](#minimax-h3-multi-reference-caption) does all of them at once — or [Universal Writer](#minimax-h3-universal-writer) describes them and writes the prompt in the same node, with their order a widget you can drag rather than a consequence of which slot you happened to use. The 8B rewriter needs none of that for its reference *frames*: connect the picture and it reads it. The Omni rewriter needs none of it at all — a clip and a sound reach it as themselves. And when the shortest way to a good prompt is somebody else's: [Prompt Presets](#minimax-h3-prompt-presets) hands over one of a thousand finished MiniMax-H3 prompts that ship inside the pack, narrowed by look and subject, each with the frame of the clip it was written for and that clip a click away. No model is loaded for it and nothing is downloaded. If your card has 8 GB, skip to [the writer nodes](#minimax-h3-prompt-writer-t2vai2vafl2val2va). ## Contents - [What you need before installing](#what-you-need-before-installing) - [Install](#install) - [Example workflows](#example-workflows) - [Nodes](#nodes) - [MiniMax-H3 Prompt Rewriter](#minimax-h3-prompt-rewriter) - [MiniMax-H3 Prompt Rewriter 8B (sees frames)](#minimax-h3-prompt-rewriter-8b-sees-frames) - [MiniMax-H3 Prompt Rewriter Omni (sees and hears)](#minimax-h3-prompt-rewriter-omni-sees-and-hears) - [MiniMax-H3 Universal Rewriter](#minimax-h3-universal-rewriter) - [MiniMax-H3 Rewriter Options](#minimax-h3-rewriter-options) - [MiniMax-H3 Prompt Writer (T2VA/I2VA/FL2VA/L2VA)](#minimax-h3-prompt-writer-t2vai2vafl2val2va) - [MiniMax-H3 Prompt Writer (Ref2VA)](#minimax-h3-prompt-writer-ref2va) - [MiniMax-H3 Universal Writer](#minimax-h3-universal-writer) - [MiniMax-H3 Reference Caption](#minimax-h3-reference-caption) - [MiniMax-H3 Multi Reference Caption](#minimax-h3-multi-reference-caption) - [Captioning with a model ComfyUI already has loaded](#captioning-with-a-model-comfyui-already-has-loaded) - [The captioner is loaded once, not once per reference](#the-captioner-is-loaded-once-not-once-per-reference) - [MiniMax-H3 Guide Prompt (any LLM)](#minimax-h3-guide-prompt-any-llm) - [MiniMax-H3 Prompt Check](#minimax-h3-prompt-check) - [MiniMax-H3 Prompt Reducer](#minimax-h3-prompt-reducer) - [MiniMax-H3 Reduce Prompt (any LLM)](#minimax-h3-reduce-prompt-any-llm) - [MiniMax-H3 Effect Embeddings](#minimax-h3-effect-embeddings) - [MiniMax-H3 LoRA Triggers](#minimax-h3-lora-triggers) - [MiniMax-H3 Reference Adapter](#minimax-h3-reference-adapter) - [MiniMax-H3 Reference Slots](#minimax-h3-reference-slots) - [MiniMax-H3 Prompt Presets](#minimax-h3-prompt-presets) - [The duration widget](#the-duration-widget) - [Repeating the last answer](#repeating-the-last-answer) - [The answer is checked](#the-answer-is-checked) - [Acting on what it found](#acting-on-what-it-found) - [The prompt library](#the-prompt-library) - [The guides are fetched, not bundled](#the-guides-are-fetched-not-bundled) - [The model list](#the-model-list) - [Models you already pulled for Ollama](#models-you-already-pulled-for-ollama) - [Where the weights go](#where-the-weights-go) - [Using a model you already have](#using-a-model-you-already-have) - [Smaller repackings](#smaller-repackings) - [If the node says a package is missing](#if-the-node-says-a-package-is-missing) - [Smallest download without any extra install](#smallest-download-without-any-extra-install) - [GGUF — smaller still, and nothing to install](#gguf--smaller-still-and-nothing-to-install) - [Progress on the node](#progress-on-the-node) - [Environment variables](#environment-variables) - [Languages](#languages) - [Notes](#notes) - [Credits](#credits) - [Licence](#licence) ## What you need before installing The LoRA route is a 27-billion-parameter language model, not a small helper. There is no way around the following, because the adapter is bound to one specific base model. **The writer nodes have none of these requirements** — see their table below. | Resource | Requirement | |---|---| | Disk | **~52 GB** for `Qwen/Qwen3.6-27B` + **~3.5 GB** for the adapter — or **~10–16 GB** total on the GGUF route | | VRAM (`nf4`, default) | **~16 GB** | | VRAM (`int8`) | ~28 GB | | VRAM (`bfloat16`) | ~54 GB, spills into system RAM via accelerate | | VRAM (GGUF) | ~13–19 GB depending on the quant, lower still with fewer offloaded layers | | Packages | `transformers`, `peft`, `accelerate`, and `bitsandbytes` for `nf4`/`int8`. **Nothing at all on the GGUF route:** `llama-cpp-python` is used when it happens to be installed, and the official llama.cpp binaries are fetched when it is not | > **The MiniMax-H3 text encoder cannot be reused for this.** It is a different > model (Qwen3-VL-32B, vocabulary 151936) from the LoRA's base (Qwen3.6-27B, > vocabulary 248320); it contains none of the linear-attention `in_proj_*` > modules the adapter targets; it is truncated to the first 50 of 64 layers; and > it ships without `lm_head` or a final norm, so it cannot generate text at all. > It only produces hidden states for the DiT. ## Install Clone into `ComfyUI/custom_nodes/` and install the requirements into the same Python environment ComfyUI runs on: ```bash cd ComfyUI/custom_nodes git clone https://github.com/pytraveler/MiniMax-H3-Prompt-Rewriter-ComfyUI ``` For ComfyUI portable on Windows: ```bat python_embeded\python.exe -m pip install -r ComfyUI\custom_nodes\MiniMax-H3-Prompt-Rewriter-ComfyUI\requirements.txt ``` Or install it from the Comfy registry through ComfyUI-Manager. ### Example workflows Seven workflows ship with the pack and appear in ComfyUI's template browser (*Workflow → Browse Templates*) under this node pack's name once it is installed. Each is a card of its own and each runs on its own: nothing is bypassed on open, and there is no second branch to mute before pressing Run. | | Template | What it needs | |---|---|---| | 1 | **Write a prompt** — one line of an idea in, a full H3 audio-video description out | one 2.6 GB GGUF | | 2 | **Rewrite a prompt with the 27B LoRA** — the LightX2V adapter this pack is named after | one 15.7 GB GGUF | | 3 | **Write a prompt from references** — the writer describes your pictures and writes from what it saw | two GGUFs, 6 GB together | | 4 | **Ready-made prompts** — a thousand finished ones, picked in a browser with their frames | nothing at all | | 5 | **Prompt to video** — ComfyUI's text-to-video template with the writer in front of it | the MiniMax-H3 weights | | 6 | **References to video** — the same for Ref2VA, the pictures reaching the generator through the writer, in the order its strip numbers them | the MiniMax-H3 ref2va weights | | 7 | **LoRA triggers and effects** — the words an adapter answers to, and MiniMax's ten effects, put into a finished prompt | nothing at all | Only 5 and 6 load a MiniMax-H3 checkpoint. The rest end at the text, which is what most of this pack is for. Those two are ComfyUI's own gallery templates with the generator folded into a single subgraph box, so what is on screen is the prompt side plus one node — the checkpoints they name are the ones the stock templates use, and ComfyUI offers to download any that are missing. Every one carries a **Read me first** note: what it does, what it downloads, what to set before pressing Run, and where to go next. The image loaders in 3 and 6 ship **empty** on purpose — a file name saved into a template points at something your machine does not have — so choose your own pictures first. The card pictures the browser draws beside each name come from `python tools/template_cards.py`, which writes them as `.jpg` next to the workflow, where the browser looks for them. For a community take, [axiomgraph's workflows](https://github.com/axiomgraph/ComfyUIWorkflow) pair the Omni rewriter with FL2VA (GPL-3.0, and they use a few extra node packs — grab them from their repository). ## Nodes ### MiniMax-H3 Prompt Rewriter The main node. It downloads whatever is missing, loads the model, generates, and releases the VRAM again. **Outputs** | Name | Contents | |---|---| | `rewritten_prompt` | The full rewrite, ready to paste into a MiniMax-H3 text input | | `integrated_multimodal_description` | Just the shot-by-shot visual section | | `overall_soundscape` | Just the diegetic audio section | | `non_diegetic_music` | Just the score section | **Inputs** - `prompt` — the short prompt to expand. - `model` — the base model. The list holds every entry from your model list plus every Qwen3.6-27B already on disk (prefixed `on disk:`). Anything not present is downloaded on first use, resuming if interrupted. The **Model list** button edits the list in a window over the graph — see below. - `resolution` / `duration` — conditions the rewrite is composed for. Keep them equal to what you pass to MiniMax-H3, or the shot pacing will not match. The list runs `48:9` down to `9:16`; the two ultrawides are the multi-monitor case — 32:9 is two 16:9 screens side by side, 48:9 three. They condition how the shot is composed, which is all this node decides; what a generator will render at that shape is its own question. - `aspect_ratio` — **the same setting, on a socket**, and it wins while something is connected. Every writer and rewriter node in the pack has it, and it takes a `STRING` or a `COMBO` link, so the primitive that already sets ComfyUI's **Resolution Selector** can drive this from the same wire. It is there because the shape of the frame is usually decided elsewhere in the graph and spelled differently there: ComfyUI's own **Resolution Selector** calls 16:9 `16:9 (Widescreen)`, a size node says `3840x1080`, a divider says `1.78`. All three are read, and a label around the pair is read through — the number pair is what counts. A frame size within 2% of a listed ratio is called by its name, so `1376x768` (which is what Resolution Selector produces for 16:9 at 1 MP) arrives as `16:9` rather than as `43:24`; a ratio that is nothing on the list passes through as itself, which is how `2.39:1` or `5:4` gets in. Something that is not a ratio at all is refused by name rather than composed for. The `resolution` widget itself has no socket, so there is one way in and it is the one that reads what other nodes write. **The picker greys out while the wire is connected** — dimmed, nothing lit, clicks refused — because a lit square would be naming a ratio the run is not going to use. **Unplugging clears the field**: an upstream node writes its value into the widget — that is how a wire feeds a widget input at all — so the text would otherwise stay behind, and it is not inert while it sits there. Anything in `aspect_ratio` outranks the picker, including a leftover. - `quantization` — how to load an *unquantized* checkpoint: `nf4` (default, ~16 GB VRAM), `int8` (~28 GB), `bfloat16` / `float16` (~54 GB). Ignored when the checkpoint brings its own quantization. - `greedy` — on by default for deterministic output. Turn it off to sample. - `seed` - `keep_model_loaded` — **off by default.** The 27B model is released as soon as the rewrite finishes, so the same GPU can run H3 video generation next. Turn it on only when iterating on prompts back-to-back. - `options` — optional; connect a **MiniMax-H3 Rewriter Options** node. ### MiniMax-H3 Prompt Rewriter 8B (sees frames) The same idea on a much smaller model that is also multimodal. LightX2V's second adapter is trained on [Qwen3-VL-8B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct), so where the 27B has to be *told* what a reference frame contains, this one is shown the frame and writes the alignment line from what it sees. It covers four tasks rather than one, and fits on a card the 27B cannot go near. ![The 8B rewriter node in ComfyUI, set to T2VA: the options node on the left, the rewriter with its first_frame and last_frame inputs in the middle, and the finished rewrite on the right with numbered shots, (S1) and (S2) speaker ids and a [English] Hello. dialogue tag](docs/node_rewriter_8b.png) **Outputs** are the same four as the rewriter above, so the two are interchangeable downstream. **Inputs** - `prompt`, `resolution`, `duration`, `greedy`, `seed` — as above. - `model` — a Qwen3-VL-8B base, in either shape the adapter is published for. A **GGUF** entry is two files from the same conversion, the model and its projector. A **safetensors** entry is the official [Qwen3-VL-8B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct) folder, which is what the adapter was trained on and what it is published as. Entries prefixed `on disk:` are already in your model folders. **Only the 8B fits the adapter**; a Qwen3-VL of another size is refused by name and number before anything is downloaded. - `quantization` — how to load a **safetensors** base: `nf4` needs about 8 GB of VRAM, `int8` about 13, `bfloat16` about 20. Ignored for GGUF, which carries its own. - `task` — `T2VA`, `I2VA`, `FL2VA`, `L2VA`. The model's own name for these is T2AV, I2AV, FL2AV and L2AV; they are the same four tasks. - `first_frame` / `last_frame` — optional IMAGE inputs. `I2VA` reads `first_frame`, `L2VA` reads `last_frame`, `FL2VA` reads both, `T2VA` reads neither. Connect the wrong one and the node says which is missing before anything loads — which end of the clip a picture belongs to is part of what the model is told. - `keep_model_loaded` — on a **safetensors** base every task honours it: the model is loaded in ComfyUI's own process and stays there. On a **GGUF** base only `T2VA` can, because the three tasks with frames run through `llama-mtmd-cli`, a fresh process each time that takes the model with it when it exits. The node says which it did rather than ignoring the switch. - `options` — the same options node as everything else. Its `adapter` dropdown lists both LoRAs; the first entry picks whichever one matches the base model you chose, so it needs no attention. **What it costs** | | Download | VRAM | |---|---|---| | Q4_K_M base + projector + Q8_0 adapter | 4.7 + 0.7 + 0.7 GB | ~9 GB | | Q8_0 base + projector + F16 adapter | 8.1 + 0.7 + 1.3 GB | ~13 GB | | safetensors base + adapter, `nf4` | 17.5 + 2.8 GB | ~8 GB | | safetensors base + adapter, `bfloat16` | 17.5 + 2.8 GB | ~20 GB | The GGUF rows pay for the runtime as well, once per machine: `I2VA`, `FL2VA` and `L2VA` run through `llama-mtmd-cli`, so the first of them fetches the official llama.cpp build — 34 MB, or 511 MB where `llama_backend` resolves to CUDA — and an installed `llama-cpp-python` does not cover it. `T2VA` is the one task that can take the wheel instead. The GGUF route is much the smaller download and needs nothing installed. The safetensors route is the shape the adapter was published in, keeps the model resident for every task rather than only for `T2VA`, and is the one to reach for if you already have the checkpoint — but it needs `transformers` and `peft`, which the pack lists as dependencies. **What to expect of it.** All four tasks produce the trained shape: `T2VA` starts straight in on the three fields, and the other three open with the alignment sentence MiniMax-H3 itself reads. The shot markers are the visible difference the adapter makes — with `use_lora` off the same model still fills the three fields, because the contract is in the system prompt, but it stops writing `[Shot 2]` cut markers and the answer comes back about a third as long. It is an 8B, and it shows in one place: the alignment line's timestamp is sometimes formatted to three decimals instead of two, and on `FL2VA` the final picture is occasionally credited to `Shot 1` rather than the last shot. The 27B does not do this. Nothing downstream parses that line, so it costs correctness nowhere — but it is worth a glance before pasting. ### MiniMax-H3 Prompt Rewriter Omni (sees and hears) LightX2V's third adapter, and the first that listens. It is trained on [Qwen2.5-Omni-7B](https://huggingface.co/Qwen/Qwen2.5-Omni-7B) — the same model this pack's captioners already use — so a reference reaches it as the asset itself: the picture, the clip, or the sound. It is also the only one of the three that covers **Ref2AV**, the full-reference task, which answers with six fields instead of three. ![The Omni rewriter node set to REF2AV: four reference sockets down the left with a checkbox on each row, eight outputs on the right from rewritten_prompt down to non_diegetic_music, and between them a strip of three coloured squares - a blue "pic 1" over ref_0, a blue "pic 2" over ref_1 and a purple "aud 1" over ref_2 - above the line "drag to reorder - that order numbers the labels - click a square to switch it off". Below that the five tasks with REF2AV lit, six aspect-ratio rectangles drawn to proportion with 16:9 chosen, a duration of 10.0, a Russian prompt, and an on-disk Qwen2.5-Omni-7B Q8_0 with its projector. On the right the finished rewrite fills all six Ref2AV fields, naming Subject 1, Subject 2, Picture 1, Picture 2 and Audio 1 across two shots](docs/node_rewriter_omni.png) **Outputs.** Seven, and which of them fill depends on the task. The four frame tasks return the same three as the other two rewriters, so they are interchangeable downstream. `Ref2AV` returns six: `subject_definitions`, `summary`, `retention_analysis`, `detailed_description`, `overall_soundscape` and `non_diegetic_music` — the same set the [Ref2VA writer](#minimax-h3-prompt-writer-ref2va) produces, and the same meanings. `rewritten_prompt` always carries the whole answer. After them comes `references`, which is not text: the switched-on references in strip order, for [Reference Slots](#minimax-h3-reference-slots) to hand on to `MiniMaxH3ReferenceToVideo` numbered the way the prompt numbers them. **Inputs** - `prompt`, `resolution`, `greedy`, `seed` — as above. - `references` — one growing socket that takes an IMAGE, a VIDEO or an AUDIO. There is no wrong socket to plug into: what a reference is called follows from what it is. Pictures are numbered among pictures and sounds among sounds, so connecting a sound between two pictures does not renumber them. - `task` — `T2AV`, `I2AV`, `L2AV`, `FL2AV`, `REF2AV`. Only `REF2AV` takes clips and sound; the other four are written from pictures alone, and connecting a sound to one of them is refused by name before anything loads. - `duration` — **the node snaps it.** MiniMax-H3 generates on a 17n+5 frame grid at 24 fps, so most lengths do not exist: ask for 10 seconds and it is 243 frames, 10.13 s, and *that* is the number written into the turn and quoted back in the alignment line. The widget is what you meant; the line and the video agree because of the snapping. - `model` — a Qwen2.5-Omni-7B base, as a **GGUF** pair (the model and its projector) or as the official **safetensors** folder. Entries prefixed `on disk:` are already in your model folders. Two kinds of near-miss are marked rather than hidden: a Qwen2.5-Omni-**3B** is `(wrong size for the adapter)`, and a **Qwen2.5-VL-7B** — which is the same architecture string, the same 28 blocks and the same width, so the adapter *would* attach — is `(vision only, not an Omni build)`, because its projector has no audio encoder and the rewrite would be about sound that was never heard. The **Open model list** button edits the `models_omni` section — see below. - `quantization` — how to load a **safetensors** base: `nf4` about 9 GB of VRAM, `int8` about 12, `bfloat16` about 20. Ignored for GGUF. **Pick the largest your card holds** — see below. - `max_frames` — how many frames to take from a clip, spread evenly. Each frame is its own picture to the model. - `reference_layout` — the strip's state as JSON. It is a widget so the arrangement travels with the workflow and through the API; the interface draws it as squares instead. - `keep_model_loaded` — honoured on a **safetensors** base. On a **GGUF** base a task with references runs through `llama-mtmd-cli`, a fresh process that takes the model with it when it exits, and the node says so rather than ignoring the switch. - `options` — the same options node as everything else. **The strip is the ordering.** Every connected reference appears as a coloured square — blue for a picture, green for a clip, purple for a sound — showing what it will be called and which socket it came from. Drag to reorder, and that order is what numbers the labels: dragging the second picture to the front is what makes it ``. Click a square to switch it off without unplugging it. The task strip greys out a task the connected references cannot serve and says why on hover, so `FL2AV` with one picture is visibly unavailable rather than a failure two minutes later. There is deliberately **no relabelling here**, unlike the Universal Writer's strip. There, a picture can be told to stand for a subject or a clip, because the guide-driven turn names its references by hand. Here the socket settles it — and a *subject* is something this adapter **produces** in `subject_definitions`, not something the request supplies. **What it costs** | | Download | VRAM | |---|---|---| | Q4_K_M base + projector + Q8_0 adapter | 4.4 + 1.4 + 0.34 GB | ~9 GB | | Q8_0 base + projector + F16 adapter | 8.1 + 1.4 + 0.65 GB | ~13 GB | | safetensors base + adapter, `nf4` | 22.4 + 1.3 GB | ~9 GB | | safetensors base + adapter, `bfloat16` | 22.4 + 1.3 GB | ~20 GB | A GGUF base pays for the llama.cpp runtime as well, once: every task but `T2VA` carries references and therefore runs through `llama-mtmd-cli` — 34 MB, or 511 MB where `llama_backend` resolves to CUDA, and an installed `llama-cpp-python` makes no difference to it. The GGUF adapter is converted from LightX2V's own safetensors with llama.cpp's `convert_lora_to_gguf.py` and published at [pytraveler/MiniMax-H3-Prompt-Rewriter-LoRA-Omni-GGUF](https://huggingface.co/pytraveler/MiniMax-H3-Prompt-Rewriter-LoRA-Omni-GGUF). The `Q8_0` build is half the size of the `F16` and behaves the same. **Quantization buys VRAM, not speed.** Which is the opposite of the intuition — a smaller model should move fewer bytes and go faster. Measured on this adapter, on one card, same prompt, same two pictures, same 256 tokens: | | VRAM | Generation | |---|---|---| | `bfloat16` | 19.4 GB | **18.6 tok/s** | | `nf4` | 8.6 GB | 15.1 tok/s | | `int8` | 12.0 GB | 6.0 tok/s | `int8` is the worst of the three on both counts: slower than `nf4` *and* larger than it. That is not a quirk of this adapter — bitsandbytes' `load_in_8bit` is LLM.int8(), which splits every matmul into an fp16 outlier part and an int8 part and recombines them, casting the activations each time. It is a scheme for fitting a model that would not fit, and it costs what it costs. `nf4` is the better small option and dequantizes on every matmul too, which is why it does not beat `bfloat16` either. So quantize only down to what the card actually holds: with 24 GB or more, `bfloat16` is both the fastest and the most faithful. The same applies to any safetensors base in this pack, since the mechanism is bitsandbytes' rather than the model's. **And none of it applies to the GGUF route**, where llama.cpp has real quantized kernels: the same FL2AV that takes 26 seconds through `bfloat16` safetensors takes 10 through `Q4_K_M`. **Pictures are scaled before the model sees them.** LightX2V's own inference script caps a picture at 301056 pixels — 384 tokens — and a frame from a clip at 100352, and this node does the same. It is not only a saving: a 1616×1616 picture is 3249 tokens, two of them overflow an 8k context before a word of the prompt is counted, and the model is being shown a shape it never saw in training. The context is then sized to what the turn actually costs, so a `Ref2AV` with eight references widens it instead of failing. **The safetensors route shows pictures only.** ComfyUI's in-process Transformers path has no way to hand the model a sound, so a clip or a sound on that base is refused rather than silently dropped from the turn. Pick a GGUF base to use them. ### MiniMax-H3 Universal Rewriter All three prompt-rewriter LoRAs in one node, with a tab at the top choosing which one runs. [Prompt Rewriter](#minimax-h3-prompt-rewriter), [Prompt Rewriter 8B](#minimax-h3-prompt-rewriter-8b-sees-frames) and [Prompt Rewriter Omni](#minimax-h3-prompt-rewriter-omni-sees-and-hears) are unchanged and still there — nothing you have already built stops working. The three adapters are not three settings of one thing. The 27B is text: Qwen3.6-27B, one task, and a reference frame reaches it only as a sentence somebody wrote. The 8B is multimodal: Qwen3-VL-8B, four tasks, and the picture itself. The Omni is multimodal and hears as well: Qwen2.5-Omni-7B, the same four tasks and a fifth of its own. Different base, different size, different download. Which is exactly why choosing between them by hand is tedious. The prompt is the same prompt, the aspect ratio is the same aspect ratio, the duration is the same duration — so trying another adapter meant retyping all of it into a second node and then keeping them in step. ![The Universal Rewriter on its Omni tab, running Ref2VA: four reference rows down the left — first_frame, last_frame, reference_video and reference_audio, each connected and switched on — and eight outputs on the right, the three every task fills first and the four Ref2VA adds after them. Three tabs across the top with "Omni LoRA / sees, hears" lit, a task strip with Ref2VA lit, eight aspect-ratio rectangles drawn to proportion from 48:9 to 9:16 with 16:9 chosen, a duration slider at 10, the prompt "A blue whale breaching at sunset, filmed from a drone", and below them model_omni pointing at an on-disk Qwen2.5-Omni-7B Q8_0 with quantization_omni set to bfloat16, then aspect_ratio, repeat_last and three buttons: Open model list, Save the last prompt and Prompt library. On the right the six-field answer, every reference marked fully_preserved in the retention analysis](docs/node_universal_rewriter.png) *One node, three tabs. Nothing above the model rows belongs to a tab — the task, the ratio, the duration and the prompt are one set of values, whichever adapter is lit. Below them each tab holds its own model and quantization, still set to whatever you last chose, including across a save and load. The task strip is the other difference: the 27B tab lights `T2VA` alone, the 8B tab adds the three frame tasks, and the Omni tab adds `Ref2VA` on top of them, which no other tab can reach.* **So the tab carries what differs, and nothing else:** | Belongs to the tab | Shared between them | |---|---| | `model_27b` / `model_8b` / `model_omni` | `prompt`, `task`, `resolution`, `duration` | | `quantization_27b` / `quantization_8b` / `quantization_omni` | `greedy`, `seed`, `keep_model_loaded`, `bypass`, `options`, both frames, the clip and the sound | The widget the other tab uses is hidden rather than reset, so it is still set to whatever you last chose when you switch back — including across a save and load. > **No captioner on the 27B tab, deliberately.** Folding a description of a > frame into the prompt does reach that adapter, and it is not wasted — the props, > surfaces and light in it turn up in the shots, in the trained shape, with > nothing leaking into the answer. But the picture is absorbed into the scene > rather than pinned to 0.00 seconds, which is exactly what the > [LoRA's own page](https://huggingface.co/lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA) > says: T2VA is finished there and FL2VA is not. A widget for it on this node > would have looked like the frame task the 27B cannot do. When you want it, > [Reference Caption](#minimax-h3-reference-caption) writes the description and > it goes in `prompt` like any other text; when the picture has to *be* a frame, > the 8B tab is the one that was trained for it. **The task switch is shared, and the 27B tab does not touch it.** On that tab it shows `T2VA` lit with the three frame tasks greyed out, because that is the honest picture of a text-only model, and clicking does nothing at all — the value the other two tabs had is still there when you switch back. On the 8B and Omni tabs a frame task greys out until the frame it is written from is actually connected and switched on, the same way `Ref2VA` does on the Universal Writer. **`Ref2VA` is on the Omni tab, with a clip socket and a sound socket to feed it.** `reference_video` and `reference_audio` are read by that task and by nothing else — the 27B and 8B adapters have no ear, and the four frame tasks take pictures alone, so connecting a sound to `FL2VA` is refused by name rather than quietly dropped. On `Ref2VA` everything connected becomes a reference the target video reuses, in socket order: `first_frame` is ``, `last_frame` is ``, then `