MiniMax H3 prompts work best when they read less like image descriptions and more like compact production plans. Instead of asking for “a cinematic woman in a rainy city,” a strong prompt explains what the viewer sees, what changes over time, how the camera moves, what the characters say, and what the audience hears.
That approach matches the model itself. Released by MiniMax on July 31, 2026, H3 is a general-purpose multimodal video model that can interpret text, images, video, and audio as a unified context. According to the official MiniMax announcement, it can create videos with native stereo sound at up to 2K resolution and durations of up to 15 seconds. This guide explains how to turn those capabilities into clearer, more controllable prompts.
What Makes MiniMax H3 Different?
H3 is designed for more than basic text-to-video generation. Its supported workflows include text-to-video, first-frame image-to-video, first-and-last-frame generation, last-frame-to-video, reference-based creation, and video editing. The open-source system produces a 768p base result and can regenerate it at 2K using the original context, helping preserve relevant detail during the higher-resolution pass.
The model also generates synchronized audio rather than treating sound as a separate afterthought. MiniMax lists 32 kHz stereo output and stable dialogue support for 11 languages, including English, Chinese, Japanese, Korean, French, German, Spanish, and Portuguese. These features make audio direction an important part of MiniMax H3 prompts—not an optional line added at the end.
The Official Structure for MiniMax H3 Prompts
MiniMax’s prompt-writing guide recommends three core fields for text and keyframe workflows:
integrated_multimodal_description: the visual and audible timeline, including style, composition, subjects, actions, camera movement, dialogue, and synchronized sounds.
overall_soundscape: ambient noise, physical action sounds, and non-verbal human sounds across the clip.
non_diegetic_music: music heard by the audience but not by characters in the scene.
This structure separates events that happen inside the scene from soundtrack elements added for the viewer. Dialogue and a radio playing in the room belong in the main timeline because characters can hear them. A background score belongs under non_diegetic_music.
For reference-heavy generation, H3 uses a more detailed six-part structure: subject_definitions, summary, retention_analysis, detailed_description, overall_soundscape, and non_diegetic_music. Labels such as <Subject 1>, <Picture 1>, <Video 1>, and <Audio 1> must keep the same meaning throughout the prompt.
Choose the Right Generation Mode First
Before writing MiniMax H3 prompts, identify the input mode:
T2VA: Build the complete audiovisual sequence from text alone.
I2VA: Use an image as the first frame, then describe how the action develops from it.
FL2VA: Supply first and last frames, then describe the continuous path between them.
L2VA: Supply the final frame and describe a plausible sequence that arrives there.
Ref2VA: Use images, videos, or audio as references for subjects, movement, style, editing, or sound.
The distinction matters. In I2VA, the opening image anchors identity, clothing, composition, and spatial relationships. In FL2VA, the prompt should concentrate on visible intermediate changes that connect the two supplied frames. The official guide generally favors a single continuous shot for FL2VA unless the creative brief explicitly requires cuts.
How to Write a Controllable Timeline
Start the first shot by defining the medium, style, framing, subject, and environment. Then describe actions in the order they occur. Add later shots only when a cut reveals genuinely new information, such as a new viewpoint, location, or moment in time.
For later shots, use a precise cut time, such as [Shot 2] At 00:04.000. Keep every timestamp inside the selected 4–15 second duration. If a simple camera move can reveal the needed detail, use it instead of inserting another cut.
Camera instructions should be written as natural actions. MiniMax recommends identifying the movement type and, when useful, its amplitude and speed. For example: “The camera pushes in with small amplitude at slow speed toward the watch in her hand.” This is more actionable than stacking vague terms such as “dynamic camera, cinematic motion.”
Concrete descriptions also outperform abstract praise. Specify warm window light, a waist-high tracking shot, fast footsteps on wet pavement, or a low cello note. Words like “beautiful,” “epic,” and “cinematic” can support a prompt, but they should not replace observable details.
Copy-Ready MiniMax H3 Prompt Example
The following eight-second T2VA example follows the official base structure:
integrated_multimodal_description: [Shot 1] Live-action commercial style, a medium-wide shot frames a ceramic artist in a sunlit studio placing a finished blue cup on a wooden table. The camera trucks right with small amplitude at slow speed as she wipes clay from her hands and turns the cup toward the window. [Shot 2] At 00:04.000, the camera cuts to a close-up of sunlight moving across the glazed surface. The artist with a warm, clear voice (S1) says: [English] Made slowly, for mornings that matter.
overall_soundscape: Soft studio room tone continues under the scrape of pottery tools and the gentle contact of ceramic against wood. A faint breeze moves the curtain.
non_diegetic_music: Sparse acoustic-guitar notes at a moderate tempo, joined by a soft sustained string tone before fading out.
For I2VA, add the official first-frame alignment instruction before the three fields:
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
Then begin Shot 1 by preserving the person, clothing, objects, layout, and style visible in <Picture 1> before describing new motion.
Common Prompting Mistakes
The most common problem is trying to fit too many events into a short clip. A 10-second video rarely needs four locations, several costume changes, multiple lines of dialogue, and complex camera choreography. Prioritize one clear visual progression.
Other mistakes include inconsistent reference labels, timestamps beyond the requested duration, dialogue repeated in the soundscape, and music placed in the wrong field. Visible signs or captions should appear in English double quotation marks, preserving the requested text exactly. Spoken dialogue should retain its original wording and language inside the <d> tags.
Finally, do not use negative prompting as a substitute for clear direction. Define what should remain stable, what should change, and when the change happens. When using references, state whether they control identity, composition, motion, voice, or overall timing.
A Practical Workflow for Better Results
Effective MiniMax H3 prompts can be built in five steps: choose the correct mode, define the opening state, map actions to the clip duration, direct camera and sound, and check every reference label for consistency. Generate at the base resolution while testing, refine only the unclear parts, and move to the 2K workflow when the composition and timing are working.
H3’s multimodal capabilities are powerful, but prompt quality still depends on good direction. Treat each prompt as a miniature storyboard with a synchronized sound plan, and the model receives a much clearer blueprint for the final video.




