MiniMax · native 2K with sound in a single pass

MiniMax H3

An omni-modal model: text, images, video and audio all land in one context, and out comes 4 to 15 seconds of native 2K where picture and stereo sound are born together. Lines are voiced with lip-sync, and an uploaded video can be reworked or continued with a plain text instruction.

4-15 sec
Duration
2K
Resolution
9 + 3 + 3
References
from —
Credits

Video examples

What MiniMax H3 can do

Native 2K

2K is the model's own resolution, not an upscale applied after generation: the frame is composed in it from the start, so fine texture — fabric, hair, lettering on objects — stays sharp.

Omni reference 9 + 3 + 3

Up to 9 images, 3 video clips and 3 audio tracks in one request. Each reference gets its own job: the face from one photo, the jacket from another, the cutting rhythm from a video.

Sound in the same pass

Lines with lip-sync, footsteps, rain, room tone and music are generated together with the picture rather than laid under it afterwards. Sound cannot be switched off — so it is worth describing in the prompt.

Editing video with words

An uploaded video can be reworked or continued with a text instruction — swap the background, change the style, extend the scene. Independent benchmarks rate this the model's strongest side.

Three input modes

Text to video; first and/or last frame; omni reference from photos, video and audio. Frames and references are different modes and cannot be combined in one request.

Voice transfer

An audio file sets the timbre and delivery of a specific character, so the model speaks in your voice. An audio reference works only alongside an image or a video.

Pricing

Could not load pricing. Try refreshing the page.

How to get a good result

1

Describe the sound. Silence is never the default: if the prompt says nothing about audio, the model invents its own ambience. Spell out the background, footsteps, voices and any score.

2

Write dialogue verbatim and in the language you want — the model voices it with lip-sync exactly as written. One short line per 5 seconds, not a paragraph.

3

Break a longer clip into shots. For 7-15 seconds set 2-5 shots with cut timestamps; one shot runs 2-5 seconds, and anything under 1.5 seconds does not read.

4

One camera move per shot, always with amplitude and speed: a slow, small-amplitude Push In. Combining several moves in one shot produces jittery footage.

5

Give every reference one non-overlapping job: take the face and hair from the first photo, do not take the location from it. The model has no negative prompt — phrase things positively.

Frequently asked questions

How is MiniMax H3 different from Hailuo 3.0?

They are the same model: MiniMax H3 is the official name, Hailuo 3.0 is what it is called inside the Hailuo app.

Can I generate video without sound?

No — audio is always generated in the same pass as the picture and cannot be disabled in the request. If you do not need it, mute it during editing or pick a model where audio is optional.

How is 2K different from the 1080p other models offer?

For H3, 2K is the native resolution: the frame is built in it rather than enlarged after generation. Price-wise it sits at the level of WAN 2.7 at 1080p, which is markedly cheaper than the other flagships.

How many references can I upload?

Up to 9 images, up to 3 videos (2-15 seconds each, 15 seconds in total) and up to 3 audio tracks. Audio only works alongside an image or a video.

Why can't I upload frames and references at the same time?

They are two different modes. Frames (first and last) set the start and end of one continuous shot, while references set appearance, style and sound. Pick one or the other.

How are input videos and images billed?

The duration of input videos is added to the clip duration and billed at the rate of the chosen resolution. The first five images are free, the sixth and beyond cost a fixed price each. Input audio is free.

Which language should I write the prompt in?

Write in your own language — AI enhancement assembles the prompt in MiniMax's official format, and a separate button translates it into English. Character dialogue stays in its original language through the translation.

Try MiniMax H3

Native 2K with sound and lip-sync in a single pass — from — credits per clip.

© 2026 Sixio. All rights reserved.

We use cookies to operate the service, keep your session, and collect anonymous statistics. See our Privacy Policy.