Omni-modal input
One model reads text, image, video and audio as a single context.
MiniMax’s omni-modal video model — native stereo audio, top-ranked editing, one Canvas.
MiniMax H3 is MiniMax’s omni-modal video model — media call it Hailuo 3.0. One model reads text, image, video and audio together and returns a 4–15 second clip at up to 2K with native stereo audio in the same pass, plus top-ranked video editing and motion transfer.
In Renoise, MiniMax H3 runs on the Canvas next to Seedance 2.0 and Kling 3.0 Omni.
Weighing it against other tools? Compare to Hailuo
One model reads text, image, video and audio as a single context.
Sound and visuals generate together in one pass — no separate audio step.
Edit regions of a video while keeping the rest of the shot intact.
Drive a subject with reference motion for consistent, directed movement.

Drag in up to 9 images, 3 videos and 3 audio clips. MD5 dedup is automatic.

Describe the shot, motion and sound in plain text, and set 4 to 15 seconds.
Choose MiniMax H3 in the model selector and hit Generate.
Turn a product into a looping hero video for a landing page.
Open a story with animated title sequences and matched sound design.
All three run inside Renoise — pick per shot, no winner. Reach for MiniMax H3 for omni-modal editing and motion transfer; Seedance 2.0 for cinematic 21:9 and 4K; Kling 3.0 Omni for native lipsync and multi-shot.
| For your shot | MiniMax H3 (Recommended) | Seedance 2.0 | Kling 3.0 Omni |
|---|---|---|---|
| Best for | Omni-modal editing & motion transfer | Audio-native, cinematic 21:9 | Lipsync & multi-shot |
| Native audio | ✓ Stereo | ✓ | — (native lipsync) |
| Region video editing | ✓ Top-ranked | — | — |
| Motion transfer | ✓ | — | — |
| Max resolution | Up to 2K | Native 4K | 1080p |
| Clip length | 4–15s | 4–15s, plus Fast mode | 3–15s (≤10s with ref video) |
| Reference inputs | Image ×9, video ×3, audio ×3 | Image ×9, video ×3, audio ×3 | Images ×7, video ×1 |
Omni-modal video with native audio, plus more models on one Canvas.
One Renoise plan unlocks MiniMax H3, plus every available video and image model.
MiniMax H3 is developed by MiniMax. Renoise integrates it alongside Seedance 2.0, Kling 3.0 Omni, Grok Imagine Video, Nano Banana 2 / Pro, GPT Image 2 and Midjourney V7 — Renoise does not train video models itself.
Effectively yes. “Hailuo 3.0” is the common media nickname; the official name is MiniMax H3, part of MiniMax’s Hailuo H series that follows Hailuo 01 and 02.
It is omni-modal — one model reads text, image, video and audio together, generates native stereo audio in the same pass, and adds top-ranked region video editing and motion transfer.
Yes. Sound effects, ambience and dialogue are generated in the same pass as the visuals and synced to the action — there is no separate audio step.
Yes. It supports region editing, so you can refine selected parts of a clip while keeping the rest of the scene, motion and audio intact. It ranks #1 for video editing on the Artificial Analysis leaderboard.
Up to 9 images, 3 videos and 3 audio clips in one generation. Renoise re-uploads references into its asset library with automatic MD5 dedup.
No. MiniMax H3 outputs up to 2K for video. If you need native 4K video, use Seedance 2.0; 4K is also available for the image models, Nano Banana Pro and GPT Image 2.
MiniMax says the H3 weights will be open-sourced, but they had not shipped at launch. In Renoise you can use MiniMax H3 today on the Canvas without hosting anything yourself.
Only with authorization. The model blocks detectable real faces; clear a face you are authorized to use first, and the cleared face then works as a reference. No public figures, celebrities or minors.
Reach for MiniMax H3 for omni-modal editing and motion transfer, Seedance 2.0 for cinematic 21:9 and native 4K, and Kling 3.0 Omni for native lipsync and multi-shot sequences. All three share one Canvas, so you can switch per shot.
MiniMax H3 runs on Renoise’s credit-based plans, from $0.34 per video, with no separate MiniMax subscription. One plan also unlocks every other available video and image model.