MiniMax H3 Max
Prompt Generator
Describe it in one line. The MiniMax H3 Max prompt generator hands it back as the fields the model was trained to read — picture first, then what the room sounds like, then the score.
It does not render anything. It writes the prompt, checks it, and sends it next door.
Your prompt, assembled
integrated_multimodal_descriptionoverall_soundscapenon_diegetic_music
ChecksSeven rules, run before you copy
- Field order matches the official sequence
- No timestamp on [Shot 1]
- Camera line names a move but not its amplitude or speed
the camera pushes in
The grey skeleton is the real format — three fields in the base modes, six when you bring references.
0 / 7,000 characters
Opens the right tool page with everything filled in.
Your text stays in your browser until you open it in a generator yourself. We do not store drafts.
What the Prompt Generator Hands You
Left is what one line became. Right is the clip that prompt produced.
This is the whole point: not prettier words, but the shape the model is actually reading.
A woman walks down a rainy street at night.
It handed back
- integrated_multimodal_description
- [Shot 1] Wide shot of a woman in a red coat walking a wet alley at night, neon reflected in standing water, slow push-in with small amplitude.
- overall_soundscape
- Rain on metal, distant traffic, footsteps in shallow water.
- non_diegetic_music
- Low synth pad, slow pulse, sitting under the scene.
Same idea, three fields, in the order the model reads them. Nothing was invented — the picture line is your sentence with the missing decisions filled in and marked as yours to change.
The Three Fields, in Reading Order
This is MiniMax's own format, not ours. They published it as a prompt writing guide in the MiniMax H3 model repository, and it fixes the fields, their order, and what belongs in each.
- integrated_multimodal_description
The picture, shot by shot.
Inside each shot the order is composition, subject, environment, action, camera, then the sound that happens on screen.
`[Shot 1]` never carries a timestamp. Every cut from the second shot onward does: `[Shot 2] At 00:04.500, …` - overall_soundscape
What the room itself sounds like.
Ambience and foley, in time order. Anything a microphone standing in the scene would pick up.
- non_diegetic_music
The score — or `N/A`.
Style, instruments, tempo and an emotional curve. Not a track name. Two sentences is plenty; over-writing this field is how people end up fighting their own soundscape.
The order is fixed. Picture, then room, then score. Swapping them is the single most common structural mistake.
Bringing material with you — a face to keep, a clip to edit? That is the six-field form, and the builder above switches to it when you pick one of those modes. The field-by-field walkthrough lives on reference to video.
Source: MiniMax’s own prompt writing guide, published in the model repository. It is not in the API docs, which is why most guides never mention it.
Five Modes, and Which Engine Runs Each
The mode decides two things at once: how many fields your prompt needs, and which engine can run it.
Pick it first — everything else follows.
| Mode | Fields | Engine | Where to run it |
|---|---|---|---|
| Text | 3 | ⚡ H3 Max · 🎬 H3 | Text to video |
| First / last frame | 3 | ⚡ H3 Max · 🎬 H3 | Image to video |
| Reference | 6 | 🎬 H3 only | Reference to video |
| Edit a clip | 6 | 🎬 H3 only | Video to video |
| Continue a clip | 6 | 🎬 H3 only | Video to video |
Three fields versus six is not easy versus hard — it is what you brought with you.
Only words or a single still? Three. Bringing material the model has to hold on to? Six.
The bottom three modes need reference input, and MiniMax H3 Max does not take reference input yet. They run on MiniMax H3 instead — at the same rate per second, with 2K available on top.
Camera Grammar: Motion, Amplitude, Speed
Camera work goes in as plain English inside the shot, and it needs three parts before the model will move anything.
- Motionpush in
- Amplitudesmall
- Speedslow
→ slow push-in, small amplitude
slow push-in, small amplitudestatic hold throughoutmedium handheld follow, medium amplitudecinematic camera workNothing in it to execute.[Push in][Truck left]That is Hailuo 02 syntax. MiniMax H3 Max reads it as text.camera movesNo type, no amplitude, no speed — so the model picks all three.
"The camera pushes in with small amplitude at slow speed" is something a renderer can execute. "Cinematic" is not.
Say the type, say how far, say how fast. Leave any of the three out and the model decides it for you.
The picker above writes the phrase and drops it into the shot for you.
Seven Checks Before You Copy
Every prompt this page assembles is run against seven rules before you see it. Flags do not block you — they tell you what is likely to come back wrong.
Field order matches the official sequence
Picture, room, score
No timestamp on the first shot
[Shot 1]
Every cut lands inside your duration
00:09.000 in an 8-second clip is a wasted run
Total length under the API ceiling
7,000 characters
Text mode has a concrete aspect ratio
adaptive is refused for text to video
Each spoken line is in the language you tagged
[Japanese] followed by English is the classic one
The action fits the seconds you asked for
This is the slow-motion check
Why your last one came back in slow motion
Slow motion is almost never a setting. It is what you get when the action you described needs longer than the seconds you asked for — the model fits it in by slowing it down.
Two fixes, both yours: cut the action, or add seconds. The check above measures roughly how much is happening and tells you before you spend anything.
All seven run in your browser as you type. Nothing is sent anywhere.
Do You Still Need Prompt Expansion?
Before it renders, MiniMax H3 Max rewrites whatever you typed into these same three fields. That step is called prompt expansion, and it has two settings.
Balanced
DefaultAbout one second
Tidies your sentence and fills the obvious gaps.
Quality
Up to 30 seconds, before rendering starts
Writes a much richer prompt. On a 5-second clip that rewrite takes longer than everything else put together.
Both figures are the model provider's own, from its published schema.
If you filled the three fields in properly here, there is not much left for it to rewrite.
That is the practical value of this page. Not prettier prose — thirty seconds you get back, and a prompt that gives you the same result twice.
Taking It to Another Model
A MiniMax H3 Max prompt is mostly plain English, so most of it travels. Three parts do not.
The shot description, the camera phrasing, the sound and music lines. Any model that reads prose will use them.
The field names themselves. Other models do not look for `overall_soundscape` — paste the contents, drop the labels.
Shot timestamps. Most models have no concept of a cut at `00:04.500`, and some reject the whole prompt for it.
Going the other way — a Hailuo 02 prompt with `[Push in][Truck left]` in it — strip the brackets first. MiniMax H3 Max reads them as text, not as instructions.
Already Have the Clip? Go Backwards
Sometimes the shot exists and the prompt does not — a reference you liked, or something you made and want to vary.
Video to prompt reads a clip and writes the fields for it. Bring the result back here to edit it, then send it on.
Want prompts that are already written and already tested? Browse the prompt library.
How to Build a MiniMax H3 Max Prompt
Three steps, no account, nothing rendered on this page.
Pick the mode
It decides how many fields you need and which engine can run it. Text and first-frame need three; reference, edit and continue need six.
5 modesSay it once, plainly
What happens, what it sounds like, and whether there is a score. The camera picker writes the movement phrase for you.
3 boxesCopy it, or send it next door
Seven checks run as you type. The open button names the tool page your mode runs on and arrives there with everything filled in.
1 click
Nothing here costs a credit, because nothing here is rendered.
MiniMax H3 Max Prompt Generator FAQ
How do I write a prompt for MiniMax H3 Max?
The MiniMax H3 Max prompt generator on this page assembles all of that from one line, so you do not have to remember it. The format is MiniMax’s own, published as a prompt writing guide in the H3 model repository.
What are the three fields, and does the order matter?
Yes, the order matters — it is fixed, and swapping the fields is the most common structural mistake. If you are bringing reference material with you, the format switches to a longer six-field form instead.
Why did my video come out in slow motion?
Either cut the action back or ask for more seconds. The seventh check on this page measures roughly how much is happening and flags it before you spend anything.
Should I use balanced or quality prompt expansion?
And if you filled the fields in properly here, there is very little left for either setting to rewrite — that is the point of writing them yourself.
Do camera instructions like [Push in] work?
Write the movement as plain English inside the shot, with all three parts: type, amplitude and speed — for example, slow push-in with small amplitude. Leave any part out and the model chooses it for you.
How long can a prompt be?
What language should I write dialogue in?
Can I use these prompts on Veo, Seedance or Kling?
Is this free, and do I need an account?
Where is the six-field reference format explained?
Build One Now, Run It Next Door
One line in, the official field order out. The MiniMax H3 Max prompt generator is free, needs no account, and hands your finished prompt straight to the tool that renders it.
Generate free · queueWritten and maintained by the MiniMax H3 Max AI Video Generator editorial teamPublished Last updated