MiniMax Hailuo 3.0
What is MiniMax Hailuo 3.0?
MiniMax Hailuo 3.0 (officially MiniMax H3, released July 31, 2026) is MiniMax's flagship video model. It renders 2K video at 24 fps with native audio, runs 5 to 15 seconds per generation, and takes three input styles: text alone, start/end frames, or reference media. It replaces the 768p/1080p ladder of the Hailuo 2.x models with a single 2K tier sized by aspect ratio, from 2560x1440 at 16:9 up to 2944x1248 at 21:9.
Key features of Hailuo 3.0
Motion physics at 2K
House test at 2560x1440, 8 seconds, prompt: slow-motion water pouring over glass on dark slate. The churn inside the glass, the refraction through it, and the droplets beading on the stone all track convincingly. The model simplified the brief from a stack of glasses to one, and the pour briefly smears into a sheet before the splash resolves.
Native audio in the same pass
House test at 2560x1440, 8 seconds: a violinist under a stone archway at dusk. The file arrives with an audio track generated alongside the picture, no separate audio step. Framing, hand placement on the bow, and the warm lantern light against the blue-hour street all hold as the camera circles.
#1 in Video Editing on Artificial Analysis
MiniMax H3 leads the Artificial Analysis Video Editing board at Elo 1,130 (July 31, 2026), ahead of Gemini Omni Flash. Send a source video as a reference plus an edit instruction in the prompt.
21:9 cinematic output
House test at 2944x1248, 10 seconds: a dawn dolly down a rain-soaked neon street. Sign text renders legible and its color spills into the wet-asphalt reflections the way real signage does. The steam plumes rise a little too symmetrically on both sides of the frame.
Reference-to-video
Pass up to 9 subject reference images, 3 reference videos, and 3 reference audio files. The model locks the subject across generations, continues a video, or applies an instruction-based edit. References cannot be combined with start/end frames in one request.
Start and end frame control
A start frame, an end frame, or both. The end frame also works alone for last-frame-guided generation. With a frame image set, output follows the input frame instead of aspect_ratio.
Best for
Character-locked series
Subject reference images keep the same character across every video in a campaign.
Instruction-based video edits
Send the source video as a reference and describe the change: swap the product, relight the scene, change the weather.
Cinematic 21:9 spots
The 2944x1248 widescreen tier renders trailer-shaped output straight from a text prompt.
Sound-on social video
Native audio means output is publishable without an audio pass.
Start-and-end frame ads
Pin the open and the close as key art and let the model animate the move between them.
Use cases
Run a product campaign where the same mascot appears in ten videos by sending the character sheet as reference images. Edit an existing video by passing it as a reference video with an instruction prompt. Render a 21:9 teaser at 2944x1248 for a landing-page hero. Build a vertical 9:16 spot with sound for social feeds in one call. For pinned compositions, pass a start frame, an end frame, or both, and the output follows the input frame's ratio.
Limitations
Hailuo 3.0 caps at 15 seconds per generation and 24 fps, with no 4K tier; rivals like Kling 3.0 render native 4K. There is no resolution parameter: output is the single 2K tier, sized by aspect ratio. Reference media and frame images are mutually exclusive in one request. On the Artificial Analysis generation boards it sits behind Gemini Omni Flash on Text-to-Video and behind Gemini Omni Flash and Seedance 2.0 on Image-to-Video with audio.
API examples
Call Hailuo 3.0 from any language by POSTing to /v1/tasks. Full parameter docs live at docs.unifically.com/models/video/minimax/hailuo.
curl -X POST https://api.unifically.com/v1/tasks \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "hailuo/minimax-3.0",
"input": {
"prompt": "A slow dolly shot across a rain-soaked neon street, reflections rippling in the puddles",
"aspect_ratio": "21:9",
"duration": 10
}
}'
Successful submission returns a task_id. Poll GET /v1/tasks/<task_id> or set a callback_url on the request to receive the finished result.
FAQs
People also ask
Hailuo 3.0 (officially MiniMax H3, released July 31, 2026) moves to a single 2K tier at 24 fps, stretches duration to any length from 5 to 15 seconds, adds 21:9 and three more aspect ratios, generates native audio in the same pass, and accepts reference images, videos, and audio for subject-locked generation and instruction-based editing.
One 2K tier. 16:9 renders at 2560x1440, 21:9 at 2944x1248, 1:1 at 1440x1440, and the portrait ratios mirror the landscape ones. There is no resolution parameter; pick the size through aspect_ratio.
No. Reference images, videos, and audio cannot be combined with start_image_url or end_image_url in one request. Use frames for pinned compositions and references for subject or style control.
Yes. Audio generates in the same pass as the video, so output arrives with sound and no separate audio step.
On the Artificial Analysis boards (July 31, 2026), MiniMax H3 is
Yes. Hailuo 3.0 requires a prompt of 2 to 2000 characters on every request, including image-to-video and reference runs.
Any whole second from 5 to 15 per generation.
Related models
Browse all models
MiniMax Hailuo
4 models: 3.0 at 2K with native audio and references (5-15s), plus 2.0/2.3/2.3 Fast at 768p-1080p (6 or 10s).
- Text to Video
- Image to Video
- Reference to Video
Veo 3.1
4 model variants, 720p–4K, lip-synced dialogue and SFX in one call.
- Text to Video
- Image to Video
- Reference to Video
- Video to Video
