Six ad variants by Friday: how I write MiniMax H3 Max prompts now
The MiniMax H3 prompt guides are for a model H3 Max does not include. The five-part brief I use instead, six controlled variants, and the frame grid.

It is Wednesday afternoon and the client wants six openings for one hook by Friday, three horizontal and three vertical. You have done this before. You know roughly what you are going to write.
Then you paste in the prompt format from the guide you found, and what comes back is fine. Not wrong. Just soft — the beats you asked for averaged into one continuous drift, the line you wanted at 3 seconds arriving whenever.
I lost most of a day to that before I worked out what was happening, and the answer is not a technique. It is that the guides ranking for MiniMax H3 prompts are describing a model that MiniMax H3 Max does not include.
The reason your copied format is not landing
The three-field structure everyone teaches — integrated_multimodal_description,
overall_soundscape, non_diegetic_music — is the input shape for a component
called H3-Context-IR. That is the layer that reads your brief and rewrites it
into the structured form the base model was actually trained on.
MiniMax's own documentation lists H3-Context-IR as absent from MiniMax H3 Max.
So the field names are not doing anything. There is no parser waiting for them. The discipline behind the format is still worth keeping — separating what we see from what we hear from what scores it is genuinely how you should think — but you are writing prose to a model, not filling in a form that something will process.
Once I stopped expecting the format to carry the weight and started carrying it myself, the outputs got sharp.
The MiniMax H3 Max brief I write now
Five labelled parts. The labels are for me, not for the model — they make the next step mechanical, which is the whole point.
STYLE Handheld creator footage, golden hour, no colour grade,
autofocus hunts once and settles.
BEATS [0 to 1.5s] She lifts the bottle into frame and the cap cracks.
[1.5 to 3.5s] She drinks, lowers it, looks straight to camera.
[3.5 to 5s] She sets it down hard on the counter and walks out
of frame left.
CAMERA One continuous handheld shot at chest height. No cuts.
AUDIO Room tone, the cap cracking on the frame of contact, a single
swallow, the bottle landing on stone. No music, no dialogue.
LIMITS Unbranded bottle, one person, no on-screen text.Four things in there matter more than the rest.
Timing blocks. [0 to 1.5s] is not a documented instruction and nothing
parses it. It still works, reliably, because it forces you to commit to an
order and gives the model an unambiguous sequence. This is the single highest
leverage habit on this model — prompt adherence on beat-sequenced briefs is
literally what fal tuned it for.
Three beats in five seconds. That is the working density. Fewer and you get the drift problem below.
Audio written as part of the shot. Picture and a 32 kHz stereo track come out of the same pass. Leave sound out and you are not getting a silent clip, you are getting the model's guess at one. Naming the frame a sound lands on — "on the frame of contact" — works far better than I expected.
Limits as a separate line. Every unbranded, one-person, no-text constraint you do not write is a thing you will be re-running for.
Six variants, one change each
Here is the actual system, and the discipline is the part people skip.
| Variant | Change only this | What it tells you |
|---|---|---|
| 1 · Hook | The [0 to 1.5s] beat | Which opening earns the second second |
| 2 · Vertical | Ratio to 9:16, re-block for a tall frame | Whether the action survives the crop |
| 3 · Length | 5s → 8s, by adding a beat, not stretching three | Whether the idea has more in it |
| 4 · Audio | Swap the AUDIO block for a spoken line | Whether it needs a voice at all |
| 5 · Locked product | Same brief, run from a still through image to video | Whether the product stays itself |
| 6 · Close | The [3.5 to 5s] beat only | Which ending survives muted autoplay |
Everything else stays byte-identical. Rewrite the style line while also changing the ratio and you have two variables and no result. Six controlled runs teach you more than twenty enthusiastic ones.
The reason this is worth the discipline here specifically: on a model that returns in under three seconds, being rigorous is free. That is the actual shift — not cheaper clips, cheaper experiments. The text to video workspace keeps the whole brief in one field so you can edit a single block between runs, and the prompt generator will build the five-part shape for you if starting from a blank box is the bit you hate.
One honest caveat on variant 2: changing the ratio is not a metadata change. A beat blocked for 16:9 usually needs re-staging for 9:16, because whatever was in the left third is now off screen. Change the ratio and the blocking, and accept that this one variant carries two variables. The other five stay clean.
The setting that makes a three-second model take thirty
Worth knowing before you conclude the speed claims are nonsense.
On the fal route there is a parameter called prompt_expansion_mode that
rewrites your prompt before generation and hands the rewrite back in an
expanded_prompt field. The heavy setting has been reported at up to around
thirty seconds of rewriting — in front of a render that takes under three.
Two things follow. If a "fast" model felt slow, check this before you check anything else. And if you are comparing MiniMax H3 Max against another model, set expansion to the same, lightest setting on both sides — otherwise you are comparing your prompt on one model against somebody's rewrite of your prompt on the other, and the difference you measure is the rewrite.
The available settings differ between the base model and MiniMax H3 Max, and the list has already changed once, so read the live schema rather than hard-coding a value. On the MiniMax v2 API the parameter does not exist at all — no expansion, no Context-IR, your words go through as written.
The frame grid, which cost me a re-cut
This one is my favourite thing about this model family and I have not seen anyone else write it down.
We pulled 49 finished clips from MiniMax's own showcase and counted frames with
ffprobe. Thirty-nine land exactly on 17n + 5 frames at 24 fps. The other
ten are transcoding artefacts — two re-encoded to 30 fps, eight off by a frame
or two.
| You ask for | Frames you get | What you actually get |
|---|---|---|
| 5 s | 124 | 5.167 s |
| 6 s | 141 | 5.875 s |
| 7 s | 175 | 7.292 s |
| 8 s | 192 | 8.000 s |
| 9 s | 209 | 8.708 s |
| 10 s | 243 | 10.125 s |
| 12 s | 294 | 12.250 s |
| 15 s | 362 | 15.083 s |
Three things that matter on a timeline:
Eight seconds is the only request that lands on a whole second. If your cut has to hit music or a fixed slot, ask for eight. This is the one I learned the expensive way.
Some requests come back short. Ask for 6, get 5.875. Ask for 9, get 8.708. Filling a nine-second gap? Request ten and trim.
Nothing warns you. A 5.167-second clip in a five-second slot leaves four frames hanging off the end, and you will find them during the review call.
Fair warning on the source: these were measured on MiniMax H3, because the official showcase is all H3 footage. Whether MiniMax H3 Max uses the same grid is not confirmed yet — we are running the boundary values now. Treat the table as a strong prior and check your own first delivery.
Why your clip came back in slow motion
Common enough to be its own genre of complaint, and it is almost always the brief.
A thin prompt plus a long duration produces slow motion. You asked for ten seconds and handed over one event. The model has to fill the time, so it stretches that one event across all of it — which is exactly what slow motion looks like.
The fix is not adding the words "normal speed". It is more beats per second. If you genuinely want one continuous action, describe what changes inside it: the camera pushes, the light shifts, someone enters at the two-second mark.
The other cause is being on the wrong engine entirely — a step-reduced local setup produces the same floaty drift. If the render came off your own GPU rather than an API, rule that out first.
Around clip fifteen, the product stops being the product
Someone in a community thread described it better than I can: "the window becomes a door."
What is available to fix it depends on your route. Through the MiniMax v2 API, H3 Max has no reference input — reference generation is still listed as coming soon. Through fal, a reference endpoint for MiniMax H3 Max now exists. On this site, reference to video runs on MiniMax H3, which has the richer and better documented reference handling.
On the plain text-to-video path, consistency is something you engineer. Two things that work:
Lock the first frame. Get one still you are happy with, then drive every variant from it. The subject becomes a fact rather than a description, and the model is only being asked to move it.
Write a character card and paste it verbatim. Not "the same woman" — a fixed block of five checkable details, reused byte-for-byte:
SUBJECT Woman, early thirties, short dark hair pushed back, faded
rust-orange technical jacket with a torn left cuff, thin white
scar through the right eyebrow, steel carabiner on the chest
strap.Five specific details beat two paragraphs of adjectives, because you can score them. When the carabiner vanishes you know before the client does.
If you need real continuity across a sequence rather than a family resemblance, that is a reference-to-video job, and the base model is the better tool for it.
Two runs are two takes, unless you say otherwise
One bit of received wisdom worth correcting: you can reproduce a run.
``seed` is an optional input on fal's MiniMax H3 Max schema, and a random one is picked when you omit it. So the accurate framing is not "no reproducibility" — it is that two runs are two different takes unless you pin the seed.
For variant work you want both directions. Pin it when you are testing a prompt change, so the difference you see is the change and not the dice. Leave it random when you like the brief and just want more takes. Most people do the second by default and then wonder why their A/B test is noisy.
MiniMax H3 Max prompt questions I get asked
Do the three-field MiniMax H3 prompts work on MiniMax H3 Max?
Not as a format with a parser behind it — H3-Context-IR is not part of H3 Max.
Keep the discipline of separating picture, soundscape and score. Do not expect
the field names to do anything.
What is the prompt length limit? 7,000 characters, same as the base model. Length is not the goal; four timed beats beat a page of adjectives.
Why is my video in slow motion? Too few events for the duration you asked for. Add beats, not adjectives.
How do I keep the same character across clips? Lock a first frame, and reuse a verbatim five-detail character block in every prompt. For genuine continuity across a sequence, use reference to video on the base model.
Can I get the same clip twice?
Yes, pin seed. Omit it and every run is a fresh take.
How long should my first clip be? Five seconds while you are finding the idea. Eight when it has to land on a whole second. What an afternoon of that costs is worked out in whole clips rather than cents per second.
Last checked 4 September 2026. The frame-grid table is ours, measured on MiniMax H3 and flagged above as unconfirmed on MiniMax H3 Max. The expansion settings have already changed once, which is why I keep telling you to read the schema instead of trusting a list. This site is independent and not affiliated with MiniMax or fal.
Written by
Editorial desk
minimaxh3max.video
Published on the MiniMax H3 Max AI Video Generator, an independent third-party interface built on MiniMax H3 Max.


