{ "id": "e3f2b845-8f2c-4b5a-9caf-eac1029d3e7e", "revision": 0, "last_node_id": 140, "last_link_id": 282, "nodes": [ { "id": 92, "type": "SaveVideo", "pos": [ 1150, 4880 ], "size": [ 1210, 126 ], "flags": {}, "order": 22, "mode": 0, "inputs": [ { "name": "video", "type": "VIDEO", "link": 260 } ], "outputs": [ { "name": "video", "type": "VIDEO", "links": null } ], "properties": {}, "widgets_values": [ "video/MiniMax_H3", "auto", "auto" ] }, { "id": 115, "type": "ResolutionSelector", "pos": [ -1480, 5610 ], "size": [ 270, 170 ], "flags": {}, "order": 0, "mode": 0, "showAdvanced": false, "inputs": [], "outputs": [ { "name": "width", "type": "INT", "links": [ 276 ] }, { "name": "height", "type": "INT", "links": [ 277 ] } ], "title": "Resolution Selector (Size)", "properties": { "Node name for S&R": "ResolutionSelector" }, "widgets_values": [ "16:9 (Widescreen)", 0.4, 32 ], "color": "#322", "bgcolor": "#533" }, { "id": 116, "type": "MarkdownNote", "pos": [ -2030, 4850 ], "size": [ 450, 889.6875 ], "flags": {}, "order": 1, "mode": 0, "inputs": [], "outputs": [], "title": "Note: MiniMax H3", "properties": {}, "widgets_values": [ "## MiniMax H3\n\n[MiniMax H3](https://www.minimax.io/blog/minimax-h3) is MiniMax's general-purpose, omni-modal generation model. It jointly understands text, image, video, and audio, and generates video with **native stereo audio**: voice, sound effects, and music are modeled jointly in a single forward pass, not layered on afterward. Output is up to 2K resolution, 24fps, and up to about 15 seconds.\n\n## ComfyUI links\n- [ComfyUI#15224](https://github.com/Comfy-Org/ComfyUI/pull/15224)\n- [šŸ¤— Comfy-Org/MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3)\n\n## About this workflow\n\nThis template runs the **reference-to-video (ref2va)** task using the `MiniMaxH3ReferenceToVideo` node. It takes any mix of reference images, videos, and standalone audio, and weaves them into the generation to lock in a character's identity, a style, a motion, a camera move, or a voice.\n\n**Key inputs**\n\n- **ref_images / ref_videos / ref_video_audios / ref_audios**: up to 9 reference images, 3 reference videos (each may carry its own paired soundtrack), and 3 standalone reference audio clips\n- **prompt**: reference the inputs by tag, in the exact order they were connected, for example ``, `