# Audio Modes and `video_ref_edit` This guide explains when an audio input is required and what happens to the final soundtrack. ## Important: `video_1` never contains audio The `video_1` socket receives an **IMAGE batch containing video frames only**. Even when those frames were loaded from a movie file that has a soundtrack, that soundtrack is not carried through the `video_1` connection. If the source video's original soundtrack is needed, extract/load it separately and connect it to `audio_1` (or another audio input where appropriate). Recommended source-video workflow: ```text Source movie ├── video frames ──> video_1 └── extracted audio ──> audio_1 ``` ## `video_ref_edit` audio requirements `video_ref_edit` always requires `video_1`. Audio depends on `audio_mode`. ## Source-performance synchronization in `video_ref_edit` When a replacement character is generated from Picture references, restoring the original soundtrack only at final mux is not sufficient to preserve speaking or singing synchronization. LongMedia therefore treats a connected paired source soundtrack as part of the source-performance contract. For `video_ref_edit`: ```text audio_mode = auto + audio_1 connected audio_mode = preserve audio_mode = preserve_reference ``` `audio_1` is encoded onto the target AV timeline and frozen as the authoritative audio clock while the video stream is regenerated. The replacement subject is generated against the exact source performance timing, and the untouched source waveform is still restored at final output. This preserves synchronization for: - mouth articulation; - speech and singing timing; - breathing rhythm; - expression timing; - head/body performance timing coupled to the source soundtrack. `reference_only` keeps connected audio as standalone H3 audio references while H3 owns the final soundtrack. `generate` also keeps H3 in control of the final soundtrack. In `video_ref_edit + lip_sync`, `audio_1` is intentionally **not** asserted to be Video1's original soundtrack: it is an independent authoritative dub/timing source, so completely different speech or singing can drive the replacement character. | `audio_mode` | Is `audio_1` required? | Final soundtrack behavior | | --- | --- | --- | | `auto` | **No** | Connected `audio_1` is preserved/restored. If Audio1 is disconnected, LongMedia uses model-generated H3 audio; Audio2/Audio3 can still remain conditioning references. | | `preserve` | **Yes** for source-audio preservation | Restores the untouched connected source audio. In `video_ref_edit`, Audio1 is also paired natively with Video1 as its soundtrack and locked to the target AV clock. | | `generate` | No | Uses model-generated H3 audio for the final output. | | `reference_only` | Only when an audio reference is intended | Connected audio can participate as an H3 reference while the final soundtrack is model-generated. | | `preserve_reference` | **Yes** | Uses connected audio as an H3 reference and restores the untouched source track at output. | | `lip_sync` | **Yes: `audio_1`** | `audio_1` is the authoritative timing/content source for native H3 lip-sync and the untouched track is restored at output. Current lip-sync setup also requires `image_1`. | ## `auto` `auto` is the flexible default. ### With `audio_1` connected ```text video_1 = source frames audio_1 = source soundtrack audio_mode = auto ``` LongMedia preserves the attached audio and restores it as the final soundtrack. In `video_ref_edit`, the connected `audio_1` is also locked into the target AV stream as the authoritative source-performance clock, so replacement facial and mouth motion are generated against the original soundtrack timing. ### Without `audio_1` ```text video_1 = source frames audio_1 = disconnected audio_mode = auto ``` This is valid. LongMedia allows H3 to produce the output audio and decodes the generated audio stream. Therefore **`audio_1` is optional for `video_ref_edit + auto`**. Audio2/Audio3 do not implicitly become the passthrough soundtrack when Audio1 is disconnected; they remain prompt-conditioning references. ## `preserve` Use `preserve` when the original soundtrack must survive unchanged. ```text video_1 = source frames audio_1 = extracted original soundtrack audio_mode = preserve ``` The source audio is restored at output rather than being replaced by the sampled H3 audio stream. In `video_ref_edit`, `preserve` also makes that source track the authoritative target-audio timing stream; it is therefore used to preserve mouth/facial performance synchronization while the replacement identity is generated. A preserve-style mode needs an actual connected source soundtrack. If it is missing, LongMedia cannot restore audio that never entered the workflow. Therefore **connect `audio_1` for `video_ref_edit + preserve`**. ## `preserve_reference` Use this when the source audio should influence H3 conditioning and the exact source waveform should also be restored at output. In `video_ref_edit`, it also activates the same authoritative source-performance timing lock used by `auto + audio_1` and `preserve`. ```text video_1 = source frames audio_1 = source soundtrack / reference audio_mode = preserve_reference ``` This mode requires connected source audio for its intended contract. ## `generate` Use this when a new soundtrack should be produced by H3. ```text video_1 = source frames audio_mode = generate ``` An input soundtrack is not required for the final-audio contract. ## `reference_only` Use this when audio is supplied as a conditioning reference while H3 still owns the final generated soundtrack. ```text audio_1 = reference audio audio_mode = reference_only ``` The audio connection is meaningful when you actually want an audio reference. The final output remains generated rather than source-audio passthrough. ## `lip_sync` For native LongMedia H3 lip-sync: ```text image_1 = visual subject / opening image audio_1 = authoritative speech or singing performance audio_mode = lip_sync ``` `audio_1` remains native H3 audio conditioning, drives the per-clip H3 Audio Guide timing, and is restored untouched at final output. For current LongMedia lip-sync, both `image_1` and `audio_1` are required. `audio_2` and `audio_3` may also be connected as additional prompt-addressable H3 references; they do not replace Audio1 as the lip-sync clock or final passthrough track. ## Quick decision table If the source movie has important original audio: ```text Want automatic behavior? -> auto + connect audio_1 Want guaranteed untouched audio? -> preserve + connect audio_1 Want audio as ref + untouched out? -> preserve_reference + connect audio_1 Want lip-sync to that track? -> lip_sync + connect image_1 + audio_1 Want a new H3 soundtrack? -> generate Want audio only as H3 reference? -> reference_only + connect audio_1 ``` If the source movie has no useful soundtrack: ```text auto -> audio_1 may stay disconnected; H3 generates audio generate -> audio_1 may stay disconnected; H3 generates audio ``` ## `video_ref_edit`: Paired Source AV Performance For character replacement/editing, `video_1` and its original soundtrack should be treated as one source performance. With `audio_mode = auto`, `preserve`, or `preserve_reference` and `audio_1` connected, LongMedia now uses two complementary timing mechanisms: 1. `video_1 + audio_1` are sent to MiniMax H3 as one native paired `video_audio` reference block. This preserves the relationship between the source facial/body performance and its soundtrack. 2. `audio_1` is also written into the target audio stream and frozen as the generation clock. The untouched source waveform is restored at final output. For `video_ref_edit`, `duration_source = auto` follows the `video_1` timeline. This prevents a slightly shorter encoded audio/container duration from cutting off the final visual performance. The intended setup is: ```text h3_mode: video_ref_edit video_1: source performance frames image_1: replacement character / identity reference audio_1: soundtrack extracted from video_1 audio_mode: preserve (or auto / preserve_reference) ``` `video_1` remains an IMAGE batch and never carries audio by itself. Connect the extracted soundtrack separately to `audio_1`. ## `duration_source`: timeline ownership is independent `duration_source` controls **only the target timeline length**. It does not decide whether connected audio is available to H3, and it does not replace `audio_mode`. In `video_ref_edit`: | `duration_source` | Target duration | Typical use | | --- | --- | --- | | `auto` | `video_1` duration | Safe source-edit default. | | `video` | `video_1` duration | Explicitly keep the source-video horizon. | | `audio` | `audio_1` duration | Redub to a shorter/longer track; longer audio can continue the scene past Video1. | | `manual` | `manual_duration` | Force any target duration. | | `longest_input` | Longest connected video/audio input | Automatically follow the longest supplied source/reference. | Examples: ```text video_1 = 6 s audio_1 = 11 s h3_mode = video_ref_edit audio_mode = lip_sync duration_source = audio ``` The target is ~11 seconds. Video1 establishes the first ~6 seconds of scene/camera/performance reference; H3 then continues the scene while Audio1 remains the authoritative redub clock. ```text video_1 = 10 s audio_1 = 6 s duration_source = audio ``` The target uses the opening ~6 seconds of Video1 and completes on the audio-owned horizon. ```text video_1 = 6 s manual_duration = 8 s duration_source = manual ``` The target is ~8 seconds regardless of input-media lengths. Changing `duration_source` never removes `