# MiniMax-H3 Turbo 4-step LoRA — ComfyUI conversion > 2026-09-04 H3-World 首帧 I2VA(P0-P3完成并通过人审):隔离的四个正式 Advanced > 节点追加在288-291,不改变旧节点ID、schema、默认值或工作流。实现固定审计 > `Danzer1xxxxChan/H3-World`提交`8174344933fe8b9fcd1cd131db177c30226f1aaf`和 > `DANNY621/H3-World`模型revision`8883340a740842b5c40b45383e5b75df34931806`。 > `step-10000.safetensors`是rank-32纯LoRA,104/104个A/B目标与本地完整FL2VA > INT8/ConvRot底模形状兼容,直接放到 > `models/loras/minimax/H3-World/step-10000.safetensors`,不需要转换。 > 推荐模型包为[`t8star/Minimax-H3-World-Comfy`](https://huggingface.co/t8star/Minimax-H3-World-Comfy), > 已按`ComfyUI/models`的相对目录整理;其文件与上述固定上游revision逐字节一致,SHA-256为 > `DDD9187B920B1E52C2D090F4E264FD83D8D433EFC2A5B159E58883AEAF96E526`,没有转换、合并或量化。 > > 首版合同固定832×480、124帧、24fps、50步、CFG 1.0和37个动作latent,只做首帧 > I2VA。每条动作句独立Qwen编码并经Token Refiner隔离;动作文本使用镜像视频时间位置, > 50层主干使用定向FlexAttention掩码,使动作只连接自己的视频时间段。非目标比例首帧执行 > scale-to-cover后居中裁切,不拉伸。当前路线拒绝其他layout/model/attention接管者,不能与 > SLA、VSA、Sol-Attn、BlockCache或其他DiT替换直接叠加。 > > 示例工作流位于`examples/workflows/26-h3-world`。最后的视频保存不使用模型驻留进程内的 > PyAV编码器,而是把RGB帧送入隔离的单线程libx264进程、按精确时长混入AAC,并在视频、音频、 > 联合三路严格解码通过后原子发布,因此FFmpeg必须可用。固定停车场素材的`forward`和`still` > 已严格串行完成,均为832×480×124、5.167秒H.264/AAC;SHA-256分别为 > `A6AC6B7D03243AE0F8D4E3F6000503F43DE4505F9707BAF7B5B6291CC516BE28`和 > `BD198D52289FB894FE6D97CDC20BF61383EAE1F0327077ADB6D935F9EA1D59DA`。匿名机械筛查无黑帧、 > 冻结帧、非有限音频或削波,且两条不是缓存重复。用户提交的hash绑定盲测记录SHA-256为 > `DA0BC93D4DBF5118DC046023030710A88335D824594B89F2E0609640DC2DECD9`:完整观看后选择A为持续前进、 > 确认稳定、画面持平且两边声音正常;揭盲后A正是`forward`。独立分析器返回 > `p3_fixed_material_gate=PASS`,因此四节点与工作流已转为正式Advanced。 > 2026-09-03 OpenVDN MiniMax H3:三个尾部节点现为正式Advanced,按固定revision加载50层 > 混合注意力分支与DMD/Stage B adapter,并通过Comfy ModelPatcher管理额外模型。DMD路线固定8 NFE, > Stage B固定50 NFE,均固定Euler/native_flow与12/3 shift。v2接受原生H3的普通T2VA、I2VA、 > 尾帧L2VA、首尾帧FL2VA、单/多参考图、参考视频+原音轨、独立参考音频和混合任务,同时拒绝 > 旧EMA_B、SLA、VSA、Sol-Attn、BlockCache或其他Attention owner。 > Windows RTX实现使用等价分组原生SDPA,不要求FA4/Triton/Diffusers补丁。 > > 权重目录为`models/diffusion_models/OpenVDN/vdn-minimax-h3`,固定HF revision > `18be6bcc4ee72585eee322ba28b5ccac2cf85ef0`;节点运行时不下载。`10-speed`内提供9份正式工作流。 > 完整ComfyUI目录包发布在`https://huggingface.co/t8star/Vdn-Minimax-H3-Comfy`,仓库根目录可直接 > 下载到`ComfyUI/models`;模型访问与使用继续受下述MiniMax H3协议和地区限制约束。 > 上游只声明T2VA;其他模式是T8在Comfy原生条件链上的真实验证扩展。完整2688列AdaLN底模的8条 > 多模态路线已逐条复跑;新版又用FL2VA pruned INT8在320×192×39串行覆盖T2VA和相同8种入口。 > Composer按`adaln_t_table`内容SHA自动选择curve-projected Turbo;每条pruned实测均精确应用 > 104个default、259个逻辑Turbo目标和51个偏置残差(310个实际Turbo补丁),日志无`ERROR lora`, > 且通过原生H.264/AAC严格联合解码。未知curve-basis签名仍fail closed。当前结构匹配底模通过显式 > `allow_structural_base`使用,报告保持 > `base_provenance_exact=false`;不作通用16GB安全承诺,机械通过也不替代逐素材画质、音频或口型人审。 > OpenVDN源码是Apache-2.0,权重则遵循MiniMax H3 Community License;其Applicable Territory排除 > 欧盟、英国、韩国和美国。完整协议已下载到模型目录,用户必须在运行前自行确认许可适用。 > 2026-09-03 Native Masked Plan B Color Match: the two isolated starter/continuation workflows now append > `MiniMaxH3LongVideoColorMatchT8Advanced` after Output Trim and before CreateVideo. It is optional and > enabled by default. The RGB-mean-only V1 was human-rejected because B changed less but both still changed, > with A's left side obvious. V2 compares at most five preceding-tail/current-head frames, applies pooled > Reinhard Lab mean/std matching plus an 8x5 local RGB residual field, caps total per-channel pixel change at > 0.02, and fades it out within 24 frames. It changes decoded SDR RGB only; native AV latent and audio bypass > it. Suspected cuts and legacy-state/checksum/chain/canvas conflicts abstain. The design was informed by > WanAnimatePlus `auto_drift` at commit `4327a9fceda22dae545969603e91bc7d7adb0bdc` and current ComfyUI built-in > ColorTransfer behavior, while the implementation adds T8-specific spatial correction, temporal fade, > atomic state, and fail-closed validation. > > The final same-seed 960x512 (0.49152MP) V2 run strictly decoded all 124/102/102-frame source clips, both > 226-frame reviews, and all anonymous review transports. Shared latent context and the segment-zero color > reference stayed unchanged. Maximum local RGB jumps fell from `0.014492` to `0.001714` and from `0.008184` > to `0.001238`; continuation minimum free VRAM was 532/515MiB. The user then reported that left-hand A still > showed a slight jump while right-hand B was much better. Reveal mapped A to soft context and B to Plan B. > This hash-bound sample therefore accepts Plan B seam-color continuity, not complete elimination in both routes; > identity/audio quality, general 16GB safety and universal Plan B superiority remain unvalidated. > 2026-09-02 SLA Precision V2 quality correction: three append-only Advanced EXP nodes pin the current > PlagueKind v1.4.3 implementation at commit `066ada9eb2378f392cc815663f63c4eef1060b4a` under MIT. > The repaired route uses FP32 pooled routing/scores, a direct Triton sparse kernel with FP32 online > softmax, sigma-derived logical step tracking, exact language/audio protection, and dense first/last > steps. The SLA LoRA is injected as a dynamic model-only residual over the FP8 base; it is never merged > and re-quantized. Old SLA nodes and workflows remain unchanged. > > The dated workflow fixes eight NFE, shifts 12/3, 32x32 blocks and 90% requested sparsity. A real > 736x416x124 rerun observed exactly `50 Dense / 6 x 50 Sparse / 50 Dense`, 20 protected blocks, and no > kernel fallback. Its decoded video/audio hashes exactly match the pre-observability same-seed run. > The same-input/same-seed dense XFormers control also strictly decoded; both clips recovered the intended > Mandarin dialogue through ASR, measured -1 SyncNet frame at 25fps, and measured +9 frames after a fixed > 400ms delayed-video control. This pair measured about 12.38% end-to-end and 18.75% sampler savings for > Precision V2. Full-speed user blind review is still pending. Minimum free VRAM was 236MiB on the latest > Precision V2 run (211MiB previously) and 245MiB on dense, so both fail the 512MiB project gate and no > universal 16GB safety claim is made. > 2026-09-02 v1.64.0 MV Vocal Lock V3 official Ref2V correction: the current recommended workflow uses the > official `minimax_h3_ref2v_turbo_4step_v0.1_comfyui_bf16.safetensors` at strength 1.0 with four > Euler/simple steps, shifts 12/3, and 1024x768 output. A same-image, same-audio, same-failing-seed > comparison removed the persistent double-face ghosting. This supersedes the earlier attribution to > seed choice or unavoidable base-model reprojection: the failed r1-r3 route had combined a generic > LarryVrh EMA Turbo LoRA with a non-official eight-step/shift-6:3 Ref2VA schedule. > > A real 32-second/five-scene r4 run then completed 5/5 scenes and 768/768 frames at 24fps. The complete > song remained outside H3 and was muxed once after assembly. Strict video-only, audio-only, and combined > decode passed; default multithreaded video decode repeated 20 times with zero anomalies after the final > all-intra baseline packaging correction. Agent review of all scenes at 2fps found no duplicate face, > background face, persistent halo, or obvious subject-edge smear. Official SyncNet measured isolated- > vocal offsets `0/-1/0/-1/0` frames at 25fps, while a 400ms delayed-video control measured nine frames. > The 32-second mechanical and proxy-visual gates pass. The user then completed review, reported that > the 32-second result had no problem and was perfect, and explicitly removed the approximately > 90-second requirement. Final human acceptance is bound to master SHA-256 > `e833277844e6980fdeacf9bdfd5c61ffe48aefdb3e1eba6869c363777b7dd75f`. Manifest `accepted` still > means mechanically saved and contract-bound for future material; this specific master additionally > has explicit human approval. > No product node calls `/prompt` or any remote API. > 2026-09-01 fully local MV / lip-scene route: three append-only Advanced EXP nodes analyze a > complete local song on CPU, compile deterministic Ref2VA prompts, render scenes strictly serially > through the connected local H3 `MODEL`, resume accepted scenes, and mux the complete original song > once after video assembly. The implementation does not submit `/prompt` or call remote H3, LLM, > TTS, music, or video APIs. This is audio-conditioned H3 performance orchestration rather than a > phoneme-level lip solver; final lip motion, identity, acting and image quality require full-speed > human review. Use the dated workflow under `examples/workflows/24-mv-lipsync`. > 2026-08-30 FlashVSR v1.1 post-processing: three append-only Advanced EXP nodes load the > official model folder, compile an explicit quality/memory plan, and restore decoded H3 frames > at 2x or 4x while returning the exact original AUDIO object. `quality_locked` keeps the public > LCSA `2.0/3.0/11` budget. `balanced_dynamic_exp` changes only eligible interior low-motion > chunks and always protects the first, last and high-motion chunks. `memory_safe` keeps the fixed > budget and uses same-seed feathered tiles plus staged offload. The LCSA mask is dispatched to a > separately installed `spas_sage_attn` block-sparse Sage2 kernel; absence or incompatibility is an > actionable error, never a silent dense fallback. Models are checked by required structure and > loadability, not filename hash, byte size or pixel area. Official FlashVSR primarily targets 4x; > the bundled 2x workflows remain conservative experiments and cannot recover missing identity, > lip sync or source detail. > 2026-08-27 v1.52.2 canvas policy: `1920x1088` is a warning/reference area, not a hard > execution cap. Conditioning, Source AV, Long Video, Multi-Keyframe, Still Image, SPEED, > Prompt Relay resource estimation and Environment Audit allow larger 32-aligned canvases and > report user-owned VRAM/runtime/OOM risk. Existing `allow_above_reference_area` inputs remain only > for old-workflow schema compatibility and no longer gate execution. > 2026-08-29 PDD integration: the existing Advanced EXP node now prefers ComfyUI's official native > PDD FinalLayer when its runtime semantics are present, and keeps the reviewed dynamic fallback for > older cores. The converted files still require the dedicated node because they contain 258 backbone > adapters plus four custom absolute 32-head banks. The node converts those banks to the official > first-head-plus-offset padded-diff layout without using a core-version, model-hash or file-size gate. > FL2VA and Ref2VA both completed serial native 736x416x22 real renders on official ComfyUI e7051b0, > with exact 8 NFE/block 0-7 selection and strict finite H.264/AAC decode. Minimum free VRAM was > 482/633MiB, so no universal 16GiB claim is made; earlier full-length fallback results remain valid. > 2026-08-30 H3 Super low-Sigma route: use the independent identity-preserve workflow when > the official LTX Stage-2 full denoise changes faces too much. Its exact schedule is > `0.5 -> 0.412 -> 0.350 -> 0` (three Euler updates), with Dense Attention as the default. > The official parity workflow is unchanged, and H3 audio continues to bypass LTX Stage 2. > Frontend workflows are organized by purpose under `examples/workflows/`, including the new > `19-pdd-acceleration` category. Each category contains an independent `README.md` with > purpose, validated outcomes, usage guidance and explicit limitations. The same hierarchy is > mirrored into the installed `MiniMax H3 T8` user-workflow menu; dated JSON filenames and graph > contents are preserved. > 2026-08-26 SLA quality correction: the append-only Turbo/SLA Profile Router now defaults to the > corrected ordinary Alpha8 Turbo LoRA at eight NFE and 12/3 shifts. A serial 736x416x124 rerun using > the difficult close-person to aerial final-frame transition strictly decoded, but full human review > rejected the result because it entered a persistent forced scene/scale transition after about one > second. Its runtime report contains eight ordinary Turbo forwards and zero SLA calls, so that clip is > evidence of incompatible FL2VA anchors, not an SLA-kernel failure or success. The recommended workflow > now repeats a same-scale anchor by default and tells users to replace it only with a compatible final > frame. Its minimum observed free VRAM was only 418MiB, below this project's 512MiB gate. The SLA > exact profile remains four model evaluations, 6/3 shifts and 85-percent dynamic sparsity: the official > LightX2V `infer_steps=5` value denotes five sigma grid points, not five model evaluations. Released > evidence covers the BF16 checkpoint family and LightX2V's FP8 recipe, not the local INT8 ConvRot base; > the new quality-oriented exact profile therefore refuses INT8 instead of silently calling it upstream > parity. Legacy SLA nodes remain loadable for diagnostics. The user-supplied file named `124f` in the > latest report actually probes as 704x416, 22 frames and 0.9167 seconds, so its filename cannot support > a conclusion about failure after one second. > The current local 1.45.0 candidate appends a fail-closed Prompt Semantic Contract Audit as node > 190 after the complete prior 189-node prefix, then appends a read-only NFE Run Contract compiler > as node 191 without moving any prior ID. It checks only user-authored required/forbidden > phrase groups plus exact dialogue and media-tag preservation. Empty anchors ABSTAIN; a mechanical > pass still returns the original prompt until the user explicitly accepts the reviewed candidate. > The captured `turns` to `stands still` provider regression is rejected. The current source has > 133 workflows and 191 unique runtime nodes; the first 190 IDs exactly match the preceding package. > The prior `final27` ZIP is a historical snapshot from before the Prompt Budget official/local > boundary correction. `final29` is now historical because the Provider Router gained a > fail-closed diagnostic for Ollama models that return only `message.thinking` without final > `message.content`. The refreshed local-only `final30` package contains 297 entries, 132 > workflows, six Quick Start subgraphs and 190 nodes. That package is now historical: the current > source adds the run-contract compiler and corrected NFE example wiring while preserving all 190 > earlier positions. Its exact filename, size and > SHA remain in excluded local verification evidence. The complete source passes 1,140 CPU tests; > this lexical audit and thinking-only refusal are not a universal semantic > equivalence or prompt-quality claim. > One further append-only Advanced EXP node now adds exact step-boundary checkpoint/resume for the > project's `dual_clock_euler` sampler with the `native_flow` schedule. It is node 188 and defaults > to disabled/no-write. Opt-in writes atomically preserve the completed packed joint-AV Euler state, > original processed noise and latent, optional packed mask, and full sigma schedule in no-pickle > safetensors. Resume validates the exact operator-declared model/run contracts, seed, runtime patch > structure, AV layout, shifts, audio-velocity protocol and tensor digests before exposing only the > remaining sigmas. A true two-process CPU split run matches the uninterrupted control bit-for-bit. > This support deliberately excludes multistep-history samplers, ancestral/SDE RNG recovery, > third-party sampler state and interruption inside a model forward; real H3 restart media equivalence > remains unclaimed. > The clean final package audit contains 292 entries, 131 frontend workflows, six Quick Start > subgraphs and 188 unique runtime nodes. Isolated extraction preserved the exact 187-node prefix > from the prior RAVEN package and appended only the NFE resume setup. > The current local append-only RAVEN integration adds nodes 185-187 without changing the first > 184 IDs or stable sampling. It does not reimplement RAVEN. The Profile node emits one exact > parameter set to both the audit and the separately installed external sampler; the Guarded Loader > checks plugin/model/CUDA/BF16 and the reviewed memory envelope before delegating; the Request > Audit calls the external runtime's own T2VA conditioning, empty-latent and causal-MODEL contracts. > The current 16GB GPU / 128GB host is deliberately outside the default gate, so no local real > RAVEN generation, quality result or OOM-safety claim exists. > The six Quick Start subgraph files use dated ASCII filenames while keeping their bilingual graph > titles and NOTE text. This avoids a reproduced Registry-pack omission of Unicode paths on a > Windows/GBK console; it does not change any subgraph graph, node schema, default, or saved stable > workflow. > Project v1.45.0 appends twenty-one opt-in planning/report/concat/provider/checkpoint nodes after all 163 v1.44 IDs: Audio > Integrity Audit, Audio Perceptual Drift Audit, Speaker Routing Audit, and Prompt Budget + Role Compiler. They never edit audio, > reassign a speaker, or truncate a prompt; the prompt compiler now preserves leading/trailing whitespace too. Risk findings return `ABSTAIN`; exact tokenizer counts > are labelled exact only when the connected CLIP exposes countable token IDs. Four dated > importable workflows include NOTE guidance. Real H3 threshold/listening calibration remains open, > so these nodes do not claim to repair model leakage, pops, tail wrap, or voice swapping. > A real CPU-only Qwen3-VL 8B/Boogu tokenizer run then compiled mixed Chinese/English text, > ``, `