--- name: minimax-h3 description: Use when writing or debugging prompts for MiniMax H3 (Hailuo 3) video-with-audio generation, running the open weights locally in ComfyUI, choosing a quant or an acceleration LoRA for the VRAM you have, wiring reference-to-video with images, video or audio, or when a generated clip produces gibberish speech, drifts off a reference identity, garbles audio after a latent upscale, or refuses to run on a build that looks current. --- # MiniMax H3 (Hailuo 3) H3 generates **video and synchronised stereo audio jointly**, from text, images, reference video, reference audio, or any mix. It is the same model family whether you hit the hosted API or run the open weights, but the two paths are wired completely differently and only one of them is free. **Two paths, do not confuse them.** - **Hosted API**, partner nodes `MinimaxHailuo03TextToVideoNode` / `...FirstLastFrameNode` / `...ReferenceNode`, category `partner/video/MiniMax`. 2K output, priced per second, no local weights. - **Local open weights**, core nodes `MiniMaxH3ImageToVideo` and `MiniMaxH3ReferenceToVideo` from `comfy_extras/nodes_minimax_h3.py`. 768p, free, and everything below is about this path. **Who owns what, so you open one file, not three.** THIS file owns the prompt format and the operating rules. `reference.md` next to it owns weights, quant sizes, acceleration packs and their wiring. The MiniMax entry in the kit's `MODELS.md` owns the node-level graph (every node and socket) and the licence. When they disagree, the node code wins and the discrepancy is a bug worth reporting. ## The prompt format is not free-form prose H3-Base consumes the output of a hosted prompt refiner (**H3-Context-IR**) that is **not** in the open release. So locally you write in the refiner's output shape yourself. From MiniMax's own `VIDEO_PROMPT_WRITING_GUIDE_base_en.md`, that is an optional instruction line, a blank line, then three fields: ``` integrated_multimodal_description: [Shot 1] ... [Shot 2] At 00:04.500, ... overall_soundscape: ... non_diegetic_music: ... ``` - `integrated_multimodal_description` carries visuals, action, shots, speakers, dialogue and diegetic sound along the timeline. `overall_soundscape` sums ambience and physical-action sound. `non_diegetic_music` is score the characters cannot hear. - **Image modes need a fixed first line.** I2V: `For the target video, at 0.00 seconds into the target video, (from [Shot 1]) is fully referenced.` First-and-last-frame uses the alignment sentence naming both pictures and the second each lands on, to two decimals. ### A complete prompt, end to end Nothing above is usable until you have seen one whole. This is a 5 s text-to-video brief in the official shape: ``` integrated_multimodal_description: [Shot 1] Cinematic medium-wide shot, Push In slowly. A bicycle mechanic in a navy work coat lowers a metal shutter in a narrow workshop at dusk; warm tungsten light spills across scattered tools and rain-dark pavement outside. He pauses, looks toward the street. At 00:03.200 he switches off the bench lamp and the frame drops to ambient blue. The mechanic (weathered voice, mid-fifties, speaks English only) says quietly: [English] That's enough for today. overall_soundscape: Steady rain on a metal awning, the rolling clatter of the shutter, one soft click of the lamp switch, distant tyres on wet asphalt. No music from within the scene. non_diegetic_music: Sparse solo piano, slow, minor key, entering after the shutter closes and fading to silence on the lamp click. ``` For **image-to-video** the same block is preceded by the fixed line and one blank line: ``` For the target video, at 0.00 seconds into the target video, (from [Shot 1]) is fully referenced. integrated_multimodal_description: [Shot 1] ... ``` Note what is doing the work: one physical action the camera and the sound can both follow, a camera move from the fixed vocabulary, the spoken line wrapped in `` so the words are exact, and the two audio fields kept separate so diegetic sound and score do not fight. ## Dialogue: the single most common cause of "the speech is gibberish" Speakers get stable IDs `(S1)`, `(S2)`, joint `(S1,S2)`. **The words go inside `` with a language tag**, and everything about who says it and how stays outside. Copy the line verbatim, do not paraphrase or translate: ``` The young woman with a quiet, breathy voice (S1) says: [English] I get off at the next station. ``` Without `` the model is never told the exact words and improvises phonetics. Voiceover needs the exact phrase `says in an off-screen voiceover` plus a statement that the lips stay closed. Use `` when a line crosses a cut, `` when speech is truncated by the end. On-screen text goes in double quotes, verbatim. ## Camera is a controlled vocabulary Motion type plus amplitude plus speed, and medium amplitude at normal speed is the default you simply omit: `Zoom In/Out`, `Push In/Pull Out`, `Pan Left/Right`, `Truck Left/Right`, `Tilt Up/Down`, `Pedestal Up/Down`, `Arc Shot`, `Tracking Shot`. ## Reference-to-video: label every file with a job `` is the one that does the real work, because it binds sources: "`` is the woman whose appearance comes from `` and whose walking motion comes from `