MiniMax H3 · Open general-purpose multimodal video model

MiniMax H3 AI Video Model

MiniMax H3 is MiniMax’s open general-purpose multimodal video model. It understands text, images, video, and audio in one context, supporting 768P/2K, 24 fps, 4–15 second video generation, reference generation, and video editing.

MiniMax H3 Ready to create

Supports text-to-video, image-to-video, first/last-frame control, and multimodal reference generation.

Prompt examples
Model overview

What is MiniMax H3?

MiniMax H3 is MiniMax’s open general-purpose multimodal video model. It places text, images, video, and audio in one context, supporting text-to-video, image-to-video, first/last-frame control, reference generation, and video editing.

According to MiniMax’s official materials, H3 generates 768P or up to 2K video at 24 fps for 4–15 seconds, with native stereo sound. It is designed for commercial content creation across advertising, branding, e-commerce, product design, UI/UX, gaming, and more.

With image, video, and audio references, H3 can understand relationships between a subject, motion, camera, style, and voice. Describe how the references relate to the target shot in natural language, and let one generation handle the multimodal context.

API model IDMiniMax-H3
Model typeOpen general-purpose multimodal video model
Output768P / 2K · 24 fps
Clip length4–15s, whole seconds
How it works

How to generate videos with MiniMax H3

Start with a prompt, first/last frame, or reference media, then create and refine a multimodal video.

01

Choose a generation mode

Choose text-to-video, first/last-frame image-to-video, or reference generation for the task at hand.

02

Add reference media

Add a first or last frame, or upload image, video, and audio references; mixed input supports up to 12 files.

03

Write the prompt

Describe the subject, camera, action, sound, and the relationship between the references and the target shot.

04

Generate and refine

Submit the asynchronous task and review the result, then keep refining with video editing or regeneration.

Official technical specifications

Built for multimodal video creation

Native stereoAudio

Video generation can produce native stereo audio with the picture

768P / 2KResolution

Supports 768P and up to 2K output at 24 fps

4–15sClip length

Set the duration in whole seconds for each generation

Open modelModel type

MiniMax describes H3 as an open general-purpose multimodal video model

Text + image + video + audioUnified context

Understand multiple input modalities and their relationships in one task

9 + 3 + 3Reference inputs

Up to 9 images, 3 videos, and 3 audio clips; mixed input is capped at 12 files

First + lastFrame control

Use 0, 1, or 2 images to control the opening and ending frames

7,000Prompt length

The API prompt limit is 7,000 characters

Core capabilities

One model for more video-making tasks

01Native stereo sound

Generate video and sound together

H3 puts video and native stereo audio in the same generation flow, making it useful for complete shots with dialogue, sound effects, music, and ambience.

Start creating
02Unified multimodal context

Understand text, images, video, and audio together

Describe relationships between references in natural language. H3 can condition on subject, motion, camera, style, and voice for multimodal reference generation and video editing.

Start creating
03Commercial content creation

From advertising to product design

MiniMax’s official materials highlight instruction following, accurate text and brand rendering, and V2V motion transfer for advertising, branding, e-commerce, product design, UI/UX, gaming, and more.

Start creating
Generation modes

MiniMax H3 generation modes

According to MiniMax’s official API documentation, H3 creates and edits video from text, first/last frames, and multimodal reference media.

ModeInputTypical use
Text-to-videoPromptGenerate video from a text description
First/last-frame image-to-videoPrompt + first and/or last frame imageControl the opening or ending frame and bring a still image to life
Reference generationPrompt + reference image, video, or audioReference a subject, motion, camera, style, voice, or editing rhythm
Video editing and regenerationPrompt + existing videoEdit a video or regenerate a qualifying 768P result at 2K
API

MiniMax H3 input limits

Reference images≤9 · 256–5760 px · ≤30 MB each
Reference videos≤3 · 2–15s each · ≤15s total · ≤50 MB each
Reference audio≤3 · 2–15s each · ≤15s total · ≤15 MB each
Mixed references≤12 image/video/audio files; audio must accompany image or video
First / last frame0–2 images · 256–5760 px · aspect ratio 2:5–5:2
Input formatsVideo H.264/H.265 · image JPG/JPEG/PNG/WEBP/HEIC/HEIF
Audio formatsWAV, MP3
Request body≤64 MB · prompt ≤7,000 characters
Prompt examples

Prompts that move

Use these prompts as starting points for cinematic clips. The locally hosted demo videos are illustrative stock footage; the prompts can guide advertising, branding, and narrative work.

City Rail in Motion

Prompt

A commuter train glides through an elevated city track, cinematic telephoto shot.

Sunset Over Open Water

Prompt

A flock of seabirds crosses a calm ocean at sunset, golden reflections on the water.

Rainy Mountain Drive

Prompt

A car follows a winding mountain road beneath heavy clouds, quiet cinematic travel footage.

Through the Wormhole

Prompt

A luminous portal opens in an abstract tunnel, light rushing toward the camera.

A Quiet Moment

Prompt

A woman pauses outdoors at dusk, soft bokeh and a contemplative mood.

Venice After Dark

Prompt

A slow boat glides through a historic canal at night, warm city lights reflected in the water.

FAQ

MiniMax H3 frequently asked questions

What is MiniMax H3?

MiniMax H3 is MiniMax’s open general-purpose multimodal video model. It understands text, image, video, and audio inputs in a unified context and supports video generation, reference generation, and video editing.

Which inputs does MiniMax H3 support?

H3 supports text prompts, images, video, and audio for text-to-video, first/last-frame image-to-video, and multimodal reference generation.

How long and how large can MiniMax H3 videos be?

The official API documentation lists 768P/2K output at 24 fps, with a duration of 4–15 seconds per generation specified in whole seconds.

Does H3 generate sound?

Yes. MiniMax describes H3 as a multimodal video model that generates video with native stereo sound in the same generation flow.

What is MiniMax H3 reference generation?

Reference generation uses reference images, videos, or audio together with a prompt. It helps the model understand a subject, motion, camera, style, voice, or editing rhythm and create the target video.

How many reference files does MiniMax H3 support?

You can use up to 9 reference images, 3 reference videos, and 3 reference audio clips, with mixed input capped at 12 files. Audio references must be accompanied by an image or video input.

Is MiniMax H3 open source?

MiniMax describes H3 as an open general-purpose multimodal video model and provides open weights. Open weights do not necessarily mean that the training data, all code, and every component are released under open-source licenses.

How do I call the MiniMax H3 API?

Submit an asynchronous video generation task with model ID MiniMax-H3, poll the task using its task_id, and retrieve the finished video from content.url when the task succeeds.

MiniMax H3

Turn a multimodal idea into usable video.

Describe an idea with text, images, video, or audio. MiniMax H3 creates 768P/2K video with native stereo sound, ready to edit and refine.

Try MiniMax H3 now

Official sources and references

The specifications and capability descriptions on this page are based on MiniMax’s official H3 launch article, video API guide, and model overview.

Example clips use locally hosted free stock footage from Mixkit. This is an independent resource and is not affiliated with MiniMax.

Specifications checked August 4, 2026