Back to Home

MiniMax H3 Hits ComfyUI: Open 2K Video With Sound

MiniMax H3 dropped today with open weights, and it is already natively supported in ComfyUI. Day zero. Not "coming soon," not "in a few weeks" — the ComfyUI team shipped native support the very morning the weights went public, and the AI video world is losing its collective mind.

This is the open-weights video moment we have been waiting for. MiniMax's third-generation video model — the successor to Hailuo 01 and Hailuo 02 — is the first in the family to go fully open, and it lands with a spec sheet that sounds too good to be true: real stereo sound generated in the same pass as the picture, output up to 2K, clips up to 15 seconds long, and one model that ingests text, images, video, and audio all at once.

H3 Is the First Open Model to Top a Video Ranking

Here is the stat that should make you sit up: according to Artificial Analysis, H3 is the first open-weight model ever to take the number one spot in a video ranking. It currently ranks first in Video Editing, second in Text-to-Video, and third in Image-to-Video — right up there with the closed giants that charge per generation.

Under the hood it is a 33-billion-parameter omni-modal powerhouse. You can throw a prompt at it that references up to nine images, three video clips, and three audio clips at once, and H3 resolves how they relate to each other — the cross-modal work happens inside the model, not across five separate tools.

What does that feel like in practice? It collapses what used to be a five-tool pipeline into a single graph.

What H3 Brings to Your ComfyUI Graph

  • First-and-last-frame control — lock the opening frame, the closing frame, or both, and let the model invent everything in between.
  • Reference-to-video — carry a subject, a motion, or even a voice through your clip using reference images, video, or audio.
  • Motion transfer — borrow a camera move, a performance, or a cutting rhythm from one clip while the subject and style come from somewhere else.
  • Native stereo audio — as the ComfyUI team puts it, audio is a property of the model, not a post-process bolted on afterward.

For graph builders, motion transfer is the sleeper hit. Combined with in-place editing, it means you can iterate on a shot the way editors have always wanted to — keep the movement, swap everything else.

Day-Zero Support, Optimized to Run on a 3060

Here is the part that should genuinely blow your mind: this thing runs locally on consumer hardware. The ComfyUI team did serious machine-learning engineering to make H3 practical outside a data center, and the numbers are wild.

  • Modulation weights — roughly 40% of the model's total parameters — were pruned and replaced with a functionally equivalent lookup table, shrinking the memory footprint with zero loss in output quality.
  • The weights ship with accurate, efficient int8 quantization plus custom kernels that cut peak VRAM during inference.
  • Total memory footprint drops 66%: from 123.6 GB in full precision down to 42.5 GB for the smallest model variants.
  • Combined with ComfyUI's dynamic VRAM offloading, that means a next-generation 2K video model running locally on a GPU like the RTX 3060.

Let that sink in. A 2K, stereo-sound, omni-modal video generator on a 3060. That is not a demo — that is a workflow you can run on your desk, free, forever.

The honest fine print: the 2K resolution module and H3-Context-IR — MiniMax's tool for translating prompts into structured intermediate formats — stay closed for now, so local output tops out at 768p and you will handle context prep yourself using MiniMax's prompting guides. And the license allows commercial use for companies under $20 million in revenue. For most creators and studios that is a non-issue; for enterprise giants it is a reason to watch where open weights are heading.

Timing adds spice: ByteDance dropped its closed Seedance 2.5 on the very same day, serving 30-second clips with built-in audio. The closed camp fired its shot, and the open camp answered with a ranking-topping, locally-runnable model — all inside the same 24 hours. When the history of this moment is written, today will be the day open video stopped being the underdog.

This changes everything for local creators. No API bills, no rate limits, no compromise. Just an open, fine-tunable, stereo-sounding video model that lives on your own GPU. The ComfyUI team called it: day zero. And day zero has never looked this good.

Comments

No comments yet. Be the first to share your thoughts!