Choose a generation mode
Choose text-to-video, first/last-frame image-to-video, or reference generation for the task at hand.
MiniMax H3 is MiniMax’s open general-purpose multimodal video model. It understands text, images, video, and audio in one context, supporting 768P/2K, 24 fps, 4–15 second video generation, reference generation, and video editing.
Supports text-to-video, image-to-video, first/last-frame control, and multimodal reference generation.
MiniMax H3 is MiniMax’s open general-purpose multimodal video model. It places text, images, video, and audio in one context, supporting text-to-video, image-to-video, first/last-frame control, reference generation, and video editing.
According to MiniMax’s official materials, H3 generates 768P or up to 2K video at 24 fps for 4–15 seconds, with native stereo sound. It is designed for commercial content creation across advertising, branding, e-commerce, product design, UI/UX, gaming, and more.
With image, video, and audio references, H3 can understand relationships between a subject, motion, camera, style, and voice. Describe how the references relate to the target shot in natural language, and let one generation handle the multimodal context.
Start with a prompt, first/last frame, or reference media, then create and refine a multimodal video.
Choose text-to-video, first/last-frame image-to-video, or reference generation for the task at hand.
Add a first or last frame, or upload image, video, and audio references; mixed input supports up to 12 files.
Describe the subject, camera, action, sound, and the relationship between the references and the target shot.
Submit the asynchronous task and review the result, then keep refining with video editing or regeneration.
Video generation can produce native stereo audio with the picture
Supports 768P and up to 2K output at 24 fps
Set the duration in whole seconds for each generation
MiniMax describes H3 as an open general-purpose multimodal video model
Understand multiple input modalities and their relationships in one task
Up to 9 images, 3 videos, and 3 audio clips; mixed input is capped at 12 files
Use 0, 1, or 2 images to control the opening and ending frames
The API prompt limit is 7,000 characters
H3 puts video and native stereo audio in the same generation flow, making it useful for complete shots with dialogue, sound effects, music, and ambience.
Start creatingDescribe relationships between references in natural language. H3 can condition on subject, motion, camera, style, and voice for multimodal reference generation and video editing.
Start creatingMiniMax’s official materials highlight instruction following, accurate text and brand rendering, and V2V motion transfer for advertising, branding, e-commerce, product design, UI/UX, gaming, and more.
Start creatingAccording to MiniMax’s official API documentation, H3 creates and edits video from text, first/last frames, and multimodal reference media.
| Mode | Input | Typical use |
|---|---|---|
| Text-to-video | Prompt | Generate video from a text description |
| First/last-frame image-to-video | Prompt + first and/or last frame image | Control the opening or ending frame and bring a still image to life |
| Reference generation | Prompt + reference image, video, or audio | Reference a subject, motion, camera, style, voice, or editing rhythm |
| Video editing and regeneration | Prompt + existing video | Edit a video or regenerate a qualifying 768P result at 2K |
Use these prompts as starting points for cinematic clips. The locally hosted demo videos are illustrative stock footage; the prompts can guide advertising, branding, and narrative work.
A commuter train glides through an elevated city track, cinematic telephoto shot.
A flock of seabirds crosses a calm ocean at sunset, golden reflections on the water.
A car follows a winding mountain road beneath heavy clouds, quiet cinematic travel footage.
A luminous portal opens in an abstract tunnel, light rushing toward the camera.
A slow boat glides through a historic canal at night, warm city lights reflected in the water.
Copy a curated prompt and start creating multimodal AI video.
MiniMax H3 is MiniMax’s open general-purpose multimodal video model. It understands text, image, video, and audio inputs in a unified context and supports video generation, reference generation, and video editing.
H3 supports text prompts, images, video, and audio for text-to-video, first/last-frame image-to-video, and multimodal reference generation.
The official API documentation lists 768P/2K output at 24 fps, with a duration of 4–15 seconds per generation specified in whole seconds.
Yes. MiniMax describes H3 as a multimodal video model that generates video with native stereo sound in the same generation flow.
Reference generation uses reference images, videos, or audio together with a prompt. It helps the model understand a subject, motion, camera, style, voice, or editing rhythm and create the target video.
You can use up to 9 reference images, 3 reference videos, and 3 reference audio clips, with mixed input capped at 12 files. Audio references must be accompanied by an image or video input.
MiniMax describes H3 as an open general-purpose multimodal video model and provides open weights. Open weights do not necessarily mean that the training data, all code, and every component are released under open-source licenses.
Submit an asynchronous video generation task with model ID MiniMax-H3, poll the task using its task_id, and retrieve the finished video from content.url when the task succeeds.
Describe an idea with text, images, video, or audio. MiniMax H3 creates 768P/2K video with native stereo sound, ready to edit and refine.
Try MiniMax H3 nowThe specifications and capability descriptions on this page are based on MiniMax’s official H3 launch article, video API guide, and model overview.
Example clips use locally hosted free stock footage from Mixkit. This is an independent resource and is not affiliated with MiniMax.
Specifications checked August 4, 2026