News
One Node, Whole Soundstage: Native Audio and Lip-Sync with MiniMax H3
ยท RenderBob team
Until recently, sound was something you bolted onto a generated clip after the fact. MiniMax H3 collapses that chain: video with native stereo audio, voice, SFX and music, in one pass, with lip-sync.

Until recently, sound was something you bolted onto a generated clip after the fact: export the silent video, open a text-to-speech tab, dig through a library for room tone and effects, drag it all into a timeline, and nudge the waveform until the mouth stopped lying. MiniMax H3 collapses that chain. It is an omni-modal, open-weight model that generates video with native stereo audio, voice, sound effects and music modelled together in a single forward pass rather than layered on afterward, at up to 2K, 24fps, and around fifteen seconds, with lip-sync. ComfyUI shipped day-zero support, which shows up as a dual VAE (one for the video stream, one for the audio) decoded together.
The quality signal is telling. On Artificial Analysis in early August 2026, H3 ranked first for video editing with audio (1,130 Elo) and sat in the top few for video without audio, so the audio is a differentiator, not an afterthought. For motion and character work, that is a meaningful shift. A talking character, a product spot with synchronised sound, a UGC-style clip, things that used to mean a multi-tool post chain, can come out of one node, coherent on the first pass.
The ecosystem has moved fast around it. Turbo LoRAs cut the schedule to 4 or 8 steps (from around 20) and keep the audio clean; community workflows chain Motion Context generations for effectively unlimited-length lip-synced video, stitching latents losslessly rather than re-encoding clips. One reported run produced a 90-second lip-synced video on a 12GB RTX 3060 in about two hours including retries: modest hardware, real output, but clearly a burst-shaped workload.
It is demanding, and it is spiky
Long, high-resolution, audio-synced generation is exactly the heavy, occasional job that wants the right hardware on demand rather than a card that is oversized the rest of the month. Iterate local, burst the heavy pass.
It is fragile across versions
A ComfyUI 0.31.0 change moved H3 onto a new audio sampling path and degraded the sound, a higher noise floor and artifacts, until people applied a legacy-sampling compatibility fix. A model working perfectly today can regress on the next update if your environment is not pinned.
Native generative audio is a genuine capability jump for motion work, one node replacing a whole post chain. Capturing it reliably means treating it like any other production model: iterate local, burst the heavy jobs, and pin the environment so an update does not silently wreck your sound the night before delivery.
More from the blog
- Visual Dubbing Goes Mainstream: Prime Video Changes the Mouth, Not the Voice
On 9 September 2026, Prime Video launched AI lip-sync for the English dub of Maxton Hall. Human actors record the dialogue. The actors' mouths are regenerated to match.
- Suggestive Editing Arrives: Story-Aware Rough Cuts Move Into the Mainstream
From Eddie AI's story-structured assemblies at NAB 2026 to Premiere's AI Assistant and Resolve 21 search, editing tools are proposing the cut, not only cleaning the footage.