Write a prompt
Describe the subject, camera move, light and mood.

No video generated yet
Enter a prompt and click generate to start creating
Describe the subject, camera move, light and mood.
Optionally attach up to two first/last-frame images.
Pick 5, 10 or 15 seconds and one of five aspect ratios.
Review the result, revise the prompt and generate again.
4
input modalities
text · image · video · audio
2K
native model output
official model capability
24
frames per second
cinematic motion cadence
5–15s
clips with audio
site presets: 5 / 10 / 15s
Model / 01
MiniMax H3 is MiniMax’s open-weights, general-purpose multimodal video model. It accepts text, images, video and audio, then generates native 2K footage at 24 fps with dialogue, sound effects and ambience in the same pass.
Finish / 02
Hand a result directly to an editing tool, or upload a clip to enhance, swap or redirect it.
Enhance and sharpen AI video to higher resolution.
Replace a face for authorized creative or localization work.
Change a head while retaining the source motion.
Redirect motion and camera language using a reference clip.
Apply the face-swap workflow to still images.
Omni-reference / 03
MiniMax H3 can combine identity, motion, visual style and sound direction in one generation.
Native 2K video + synchronized audio
Showcase / 04
Six MiniMax H3 examples spanning brand film, mixed-media creativity, narrative, product advertising, interface motion and stylized animation.
A cinematic fashion concept that coordinates character, product, setting and camera language.
Live-action footage blends with luminous hand-drawn animation and environmental sound.
A vertical trailer builds character continuity, tension and shot-to-shot progression.
A rapid fashion edit keeps the product readable while varying faces and framing.
A product interface becomes a polished interaction demo with spatial transitions.
Bold English typography and a graphic-novel hero reveal create a compact trailer beat.
Specifications / 05
Use cases / 06
Cinematics, character PVs and animated UI concepts with stable visual identity.
Clay, pixel art, anime and 3D fantasy with a consistent look across shots.
Turn product stills into demonstrations, explainers and campaign concepts.
Short-form concepts for social, display and connected-TV campaigns.
Direct trailers, brand films and short drama with camera language and sound.
Portrait clips for Shorts, Reels and TikTok-ready storytelling.
Compare / 07
| Capability | MiniMax H3 | Hailuo 2.3 |
|---|---|---|
| Maximum resolution | Native 2K | 1080P |
| Maximum length | 15 seconds | 10 seconds |
| Native audio | Yes | No |
| Reference inputs | Images + video + audio | First frame |
| Prompt limit | 7,000 characters | 2,000 characters |
| Weights | Open | Closed API |
Prompt lab / 08
Copy a prompt and paste it into the generator above. MiniMax H3 is already selected.
A lone traveler steps from a vintage coupe on a desert highway and lifts a leather bag into the wind. Begin with a wide tracking shot, move into tactile product details, warm late-afternoon light, restrained engine noise and desert ambience.
In a quiet kitchen, a hand places a glass on the counter and glowing hand-drawn vines spread from the point of contact. Handheld phone framing, natural morning light, subtle room tone and a soft sparkling sound as the sketch animation grows.
A young vampire and a human detective meet beneath a rain-soaked station clock as a distant train arrives. Build from an uneasy two-shot to close-ups, blue moonlight and red practical lights, rain and rail ambience, ending on their locked gaze.
Create a vertical eyewear campaign with three confident models passing one pair of silver glasses between match cuts. Crisp studio lighting, quick push-ins, rhythmic shutter sounds and clean graphic backgrounds; finish on a centered product hero shot.
Turn a floating AI assistant card into a premium website interaction demo. The card expands into a layered workspace, panels snap into place with smooth parallax, cool gradients and quiet interface sounds, ending on the completed dashboard.
A sword cultivator crosses a cloud bridge while spirit cranes circle a distant mountain temple. Use a sweeping crane shot into a fast side track, luminous mist, teal-and-gold 3D fantasy styling, wind and bell ambience, ending as the temple gate opens.
Direction / 09
H3 responds best when a prompt directs change over time, not just a static image. Give each of these six elements a clear job.
Subject and intent
Name the main subject, its goal and the details that must remain recognizable.
Action over time
Describe the opening state, the key movement or change, and how the shot should end.
Scene and response
Set the location, weather and atmosphere, then explain how the environment reacts to the action.
Camera language
Choose framing and movement—wide, close-up, handheld, tracking, crane or a deliberate transition.
Look and sound
Direct lighting, palette and texture together with dialogue, ambience, effects or musical rhythm.
Reference consistency
Say what each reference controls—identity, product, composition, motion or style—and what must not change.
FAQ / 10
MiniMax H3 is a general-purpose multimodal video model from MiniMax. It can interpret text and reference media, plan motion across a shot, and generate video with synchronized dialogue, effects and ambience.
They refer to the same model family in common product usage. MiniMax H3 is the official model name; Hailuo 3 is the name many users and third-party tools use to connect it with earlier Hailuo releases.
This generator offers 5, 10 and 15-second clips. For a short clip, describe one clear action arc rather than packing in several unrelated scenes.
Yes. H3 can generate dialogue, sound effects and ambience together with the visuals. Include the speaker, exact spoken line, environment and desired sound texture in the prompt when audio matters.
This generator offers 480P, 720P and 1080P, with 16:9, 9:16, 4:3, 3:4 and 1:1 ratios. Choose 9:16 for vertical social video, 16:9 for landscape storytelling, and 1:1 when the destination needs a square asset.
Omni-reference is H3’s ability to use images, video and audio as coordinated guidance for identity, motion, visual style and sound. This page currently provides text-to-video plus first/last-frame image guidance, a focused subset of the full model.
MiniMax describes H3 as an open-weights model. Self-hosting still requires suitable GPU capacity, model files and a compatible inference stack; this page is the simpler hosted path for occasional generation.
Compared with Hailuo 2.3, H3 adds native synchronized audio, richer multimodal reference control, longer instructions, higher native resolution capability and open weights. The practical benefit is more control over a complete shot, not only a stronger first frame.
You can write a prompt and prepare reference images before signing in. Clicking Generate Video for Free opens the account dialog; new accounts receive 20 credits, enough for two default 480P, 5-second generations.
Write the subject and goal first, then describe the action in order, the environment, camera movement, lighting and sound, and the final frame. One coherent sequence usually gives the model clearer direction than a list of unrelated visual adjectives.
Use a first frame when the opening composition, character or product must be specific. Add a last frame when the shot needs to arrive at a particular composition; your prompt should explain the motion that connects the two images.
Generation time varies with clip length, resolution and queue load. After a task is accepted, its state is stored in your account history; keep the page open for live progress, or return to the history page to check the completed result.
A failed task shows a specific error whenever the provider returns one. Credits reserved for an unsuccessful task are restored automatically; if a status remains stuck, use the task ID when contacting support.
Commercial use depends on the model and service terms plus the rights to every prompt, logo, face, voice and reference asset you provide. Generation does not automatically clear copyright, trademark, likeness or voice rights.
Only use media you own or are authorized to edit. Unauthorized impersonation and deceptive use are prohibited.
Sign up, claim 20 credits and generate your first video with audio on this page.