MiniMax H3 Max Video to Prompt
Watch a clip, get the MiniMax H3 Max prompt that could make one like it — all three fields, in the order the model expects.
Your structured prompt
integrated_multimodal_descriptionshots · camera · dialogueoverall_soundscapewhat the scene makesnon_diegetic_musicor None.
Split into fields, not prose — with the timestamps already refitted.
Try a sample clip: , ,
Before you copy
- Style is fine; a recognisable work is not.
- Real, identifiable people need consent.
- Describe a mood, not a specific song.
Your first clip on this site is free — no card.
What MiniMax H3 Max video to prompt
actually extracts
A MiniMax H3 Max prompt, not a caption. What comes back is the three labelled fields the model actually reads — a numbered shot list with camera moves and timestamps, the soundscape, and the score — rather than a paragraph describing the picture. Generic tools return the paragraph, and half of what your clip is doing is audible.
Your clip
Read out as fields, not prose
MiniMax H3 Max video to prompt watches your clip — the frames and the audio — and writes back the structured prompt that could produce one like it: a numbered shot list with camera moves and timestamps, the scene's sound, and the score, in the three labelled fields H3 reads. Generic tools return a paragraph about the picture. H3 does not take paragraphs, and half of what your clip is doing is audible. Paste a link or drop a file; the first extraction each week needs no account.
- Passes over the clip
- 7
- Labelled fields out
- 3
- Source it will read
- ≤ 60 s
- Timestamps refitted to
- 4–15 s
Passes over the clip
Labelled fields out
Source it will read
Timestamps refitted to
- No keyframe sampling
- No paragraph of prose
- No account
Five things people take from a reference clip into MiniMax H3 Max
Naming the target changes which pass matters most, how exact the wording has to be, and whether a written prompt is even the right vehicle.
From footage to MiniMax H3 Max fields: the seven passes
The extraction reads a shot the way the official prompt guide orders one: composition, subject, environment, action, camera, sound — and the exact moment each referenced thing appears.
Cuts
How many shots there are and where they break.
[Shot n] · At 00:0x.xxx
Composition
Shot size and camera height, read per shot.
Opens each shot
Subject
Specific enough that a stranger could redraw it — "a woman in her early 30s, short black bob, oversized grey trench coat", never "a woman".
Subject description
Environment
Place, weather and the direction of the light. H3 treats light as physics, so "window light from frame left" beats "beautiful lighting".
Environment description
Action
What actually happens, in order.
Action description
Camera
Written the way H3 executes it: type, amplitude, speed. "Slow push-in, small amplitude" survives the trip. "Smooth cinematic movement" does not — there is nothing in it to execute.
Camera description
Sound
The three layers below, and the pass every other tool skips.
Two sound fields + dialogue
The three sounds your clip is making
Keyframe tools sample stills, which means they analyse your clip on mute. This extraction listens to the audio track and splits what it hears three ways — dialogue, soundscape and score — because MiniMax H3 Max takes each layer in a different field. The test between the last two is one question: could the people in the clip hear it?
The extraction here listens, and splits what it hears three ways, because H3 takes each layer in a different place. The test between the last two is one question — could the people in the clip hear it?

Dialogue
The words people actually say, transcribed with a speaker ID, a delivery and a language tag, landing inside the shot description.
(S1) says warmly, [English] …
Soundscape
Everything the scene itself produces: rain on an awning, a knife on a board, traffic two streets back — each source with a distance and a moment it lands.
overall_soundscape
Score
What only the audience hears: instruments, tempo, where it swells and where it stops. If the clip has none, the field says None — itself an instruction H3 obeys.
non_diegetic_music
That split is why a MiniMax H3 Max video to prompt result has two sound fields at the end, and why a result without them wastes half the model.
A 30-second clip does not fit a 15-second clock
MiniMax H3 Max generates 5 to 15 seconds and the base model takes 4, so your reference probably is not that. Copying a 30-second clip’s timestamps across unchanged breaks an explicit official rule — timing that contradicts the requested duration is rejected — so the extractor remaps the beats instead of transcribing them.
Copying its timestamps across unchanged breaks an explicit official rule — timing that contradicts the requested duration is rejected — so the extractor remaps instead. Pick a target length and the beats that matter (cuts, action peaks, sound events) are kept and re-spaced. Or drag out one section of the source and take only that; a tight four seconds usually reverses better than a loose thirty.
Two working numbers from testing: a shot with a real action change needs about three seconds, and a static mood shot survives on two — so eight seconds carries about two shots, fifteen about three. The picker shows the beats it kept and the beats it dropped, so you can veto the machine's taste.
Frames snap to H3's 17n+5 grid at 24 fps, and every remapped timestamp is checked against the duration you chose before you copy anything.
When MiniMax H3 Max video to prompt is the wrong tool
Preserve with the file, recreate with the prompt. A prompt is a description, and a description is lossy on purpose — which is what you want for learning why a clip works, and the wrong tool for keeping its exact framing, lighting or soundtrack. For those, hand the model the clip itself through reference to video or video to video.
A prompt is a description, and a description is lossy on purpose. If the goal is learning why a clip works, or making a new one in its spirit, the loss is the point. If the goal is keeping the original framing, lighting or soundtrack, stop describing and hand H3 the clip itself — reference mode preserves what a rewrite can only approximate, and re-voicing or continuing the actual footage is video to video's job.
The one case where the prompt is the only path is when the source cannot be uploaded at all — someone else's footage, private material — because a description you wrote is yours in a way a copy never is.
Recreating someone else's clip: where the line is
Style is not protectable; a specific work is. Reversing a clip to learn how a look was built, and then making something of your own in that register, is ordinary practice. Reproducing a particular film shot for shot, or a recognisable person’s face or voice without their consent, is not — and no prompt extractor changes that.
- StyleLearning that a clip leans on golden-hour light and a slow push-in is fine.
- A specific workReproducing a recognisable work shot for shot is not.
- A real personDescribing a real, identifiable person into a model is not either.
- MusicDescribe a mood and an instrumentation; do not reconstruct a specific song.
- Logos and productsProtected no matter how the video was made. If your output would be recognisable as someone else's work, ask them first.
How to turn a reference clip into a MiniMax H3 Max prompt
MiniMax H3 Max video to prompt reads frames and audio, not sampled keyframes, which is why the sound comes back as its own fields.
Paste a URL or upload a clip
Up to 60 seconds and 50 MB. A direct link works as well as a file.
≤60 s, ≤50 MBLet the seven passes run
Cuts, composition, subject, environment, action, camera and sound are read separately.
Seven passesRead the shot list and the two sound fields
Dialogue comes back transcribed with speaker ID, delivery and language tag; soundscape and score stay separate.
Official field orderTake the refitted prompt
Timestamps are remapped from your clip's length to a 4–15 second target, so the result is generatable as-is.
Remapped to 4–15 s
One extraction a week without an account, three when signed in, then 5 credits each.
MiniMax H3 Max video to prompt FAQ
What to know before you turn a clip back into a MiniMax H3 Max prompt
What is MiniMax H3 Max video to prompt?
Will I get the exact same video back?
Why not use a generic video to prompt tool?
Should I just use the clip as a reference instead?
My source is forty seconds long. What happens?
Does it transcribe the dialogue?
What do I do with the prompt?
Is it legal to recreate someone else's video?
What can I paste or upload, and how long can it be?
Does it work on clips with no dialogue?
Is it free?
Can it go the other way — idea to prompt?
Paste the clip, read it back as a prompt.
One free MiniMax H3 Max extraction a week without an account, and what you copy is yours to keep.
Generate free · queueWritten and maintained by the MiniMax H3 Max AI Video Generator editorial teamPublished Last updated