MiniMax H3 Max Video to Prompt

Watch a clip, get the MiniMax H3 Max prompt that could make one like it — all three fields, in the order the model expects.

Read a clip back

What are you after?

Target lengthH3 generates 4–15s

  • Cost1 free a week signed out, 3 signed in. Then 5 credits each.
  • Your clipRead in your browser. Never uploaded.

Your structured prompt

  • integrated_multimodal_descriptionshots · camera · dialogue
  • overall_soundscapewhat the scene makes
  • non_diegetic_musicor None.

Split into fields, not prose — with the timestamps already refitted.

Try a sample clip: , ,

Before you copy

  • Style is fine; a recognisable work is not.
  • Real, identifiable people need consent.
  • Describe a mood, not a specific song.
Responsible use

Your first clip on this site is free — no card.

What comes back

What MiniMax H3 Max video to prompt actually extracts

A MiniMax H3 Max prompt, not a caption. What comes back is the three labelled fields the model actually reads — a numbered shot list with camera moves and timestamps, the soundscape, and the score — rather than a paragraph describing the picture. Generic tools return the paragraph, and half of what your clip is doing is audible.

A reference clip being read frame by frame

Your clip

Read out as fields, not prose

MiniMax H3 Max video to prompt watches your clip — the frames and the audio — and writes back the structured prompt that could produce one like it: a numbered shot list with camera moves and timestamps, the scene's sound, and the score, in the three labelled fields H3 reads. Generic tools return a paragraph about the picture. H3 does not take paragraphs, and half of what your clip is doing is audible. Paste a link or drop a file; the first extraction each week needs no account.

Passes over the clip
7

Passes over the clip

Labelled fields out
3

Labelled fields out

Source it will read
≤ 60 s

Source it will read

Timestamps refitted to
4–15 s

Timestamps refitted to

  • No keyframe sampling
  • No paragraph of prose
  • No account

Five reasons to reverse a clip

Five things people take from a reference clip into MiniMax H3 Max

Naming the target changes which pass matters most, how exact the wording has to be, and whether a written prompt is even the right vehicle.

The method

From footage to MiniMax H3 Max fields: the seven passes

The extraction reads a shot the way the official prompt guide orders one: composition, subject, environment, action, camera, sound — and the exact moment each referenced thing appears.

  1. Pass 01Cuts

    How many shots there are and where they break.

    [Shot n] · At 00:0x.xxx

  2. Pass 02Composition

    Shot size and camera height, read per shot.

    Opens each shot

  3. Pass 03Subject

    Specific enough that a stranger could redraw it — "a woman in her early 30s, short black bob, oversized grey trench coat", never "a woman".

    Subject description

  4. Pass 04Environment

    Place, weather and the direction of the light. H3 treats light as physics, so "window light from frame left" beats "beautiful lighting".

    Environment description

  5. Pass 05Action

    What actually happens, in order.

    Action description

  6. Pass 06Camera

    Written the way H3 executes it: type, amplitude, speed. "Slow push-in, small amplitude" survives the trip. "Smooth cinematic movement" does not — there is nothing in it to execute.

    Camera description

  7. Pass 07Sound

    The three layers below, and the pass every other tool skips.

    Two sound fields + dialogue

The pass others skip

The three sounds your clip is making

Keyframe tools sample stills, which means they analyse your clip on mute. This extraction listens to the audio track and splits what it hears three ways — dialogue, soundscape and score — because MiniMax H3 Max takes each layer in a different field. The test between the last two is one question: could the people in the clip hear it?

The extraction here listens, and splits what it hears three ways, because H3 takes each layer in a different place. The test between the last two is one question — could the people in the clip hear it?

  • Two women in hanfu on a mountain terrace, mid-exchange, one caught eating

    Dialogue

    The words people actually say, transcribed with a speaker ID, a delivery and a language tag, landing inside the shot description.

    (S1) says warmly, [English] …
  • A street frozen mid-motion, debris and pedestrians held in the air while the camera keeps moving

    Soundscape

    Everything the scene itself produces: rain on an awning, a knife on a board, traffic two streets back — each source with a distance and a moment it lands.

    overall_soundscape
  • A hand-lettered title card reading thoughtcrime, a figure silhouetted at a desk beneath it

    Score

    What only the audience hears: instruments, tempo, where it swells and where it stops. If the clip has none, the field says None — itself an instruction H3 obeys.

    non_diegetic_music

That split is why a MiniMax H3 Max video to prompt result has two sound fields at the end, and why a result without them wastes half the model.

Timeline remap

A 30-second clip does not fit a 15-second clock

MiniMax H3 Max generates 5 to 15 seconds and the base model takes 4, so your reference probably is not that. Copying a 30-second clip’s timestamps across unchanged breaks an explicit official rule — timing that contradicts the requested duration is rejected — so the extractor remaps the beats instead of transcribing them.

Copying its timestamps across unchanged breaks an explicit official rule — timing that contradicts the requested duration is rejected — so the extractor remaps instead. Pick a target length and the beats that matter (cuts, action peaks, sound events) are kept and re-spaced. Or drag out one section of the source and take only that; a tight four seconds usually reverses better than a loose thirty.
Two working numbers from testing: a shot with a real action change needs about three seconds, and a static mood shot survives on two — so eight seconds carries about two shots, fifteen about three. The picker shows the beats it kept and the beats it dropped, so you can veto the machine's taste.

Frames snap to H3's 17n+5 grid at 24 fps, and every remapped timestamp is checked against the duration you chose before you copy anything.

Honest routing

When MiniMax H3 Max video to prompt is the wrong tool

Preserve with the file, recreate with the prompt. A prompt is a description, and a description is lossy on purpose — which is what you want for learning why a clip works, and the wrong tool for keeping its exact framing, lighting or soundtrack. For those, hand the model the clip itself through reference to video or video to video.

  • Learn from it

  • Make one like it

  • Keep it

  • Change it

  • Cannot upload it

A prompt is a description, and a description is lossy on purpose. If the goal is learning why a clip works, or making a new one in its spirit, the loss is the point. If the goal is keeping the original framing, lighting or soundtrack, stop describing and hand H3 the clip itself — reference mode preserves what a rewrite can only approximate, and re-voicing or continuing the actual footage is video to video's job.
The one case where the prompt is the only path is when the source cannot be uploaded at all — someone else's footage, private material — because a description you wrote is yours in a way a copy never is.

Where the line is

Recreating someone else's clip: where the line is

Style is not protectable; a specific work is. Reversing a clip to learn how a look was built, and then making something of your own in that register, is ordinary practice. Reproducing a particular film shot for shot, or a recognisable person’s face or voice without their consent, is not — and no prompt extractor changes that.

  • StyleLearning that a clip leans on golden-hour light and a slow push-in is fine.
  • A specific workReproducing a recognisable work shot for shot is not.
  • A real personDescribing a real, identifiable person into a model is not either.
  • MusicDescribe a mood and an instrumentation; do not reconstruct a specific song.
  • Logos and productsProtected no matter how the video was made. If your output would be recognisable as someone else's work, ask them first.

Full policy: Responsible use

The whole procedure

How to turn a reference clip into a MiniMax H3 Max prompt

MiniMax H3 Max video to prompt reads frames and audio, not sampled keyframes, which is why the sound comes back as its own fields.

  1. Paste a URL or upload a clip

    Up to 60 seconds and 50 MB. A direct link works as well as a file.

    ≤60 s, ≤50 MB
  2. Let the seven passes run

    Cuts, composition, subject, environment, action, camera and sound are read separately.

    Seven passes
  3. Read the shot list and the two sound fields

    Dialogue comes back transcribed with speaker ID, delivery and language tag; soundscape and score stay separate.

    Official field order
  4. Take the refitted prompt

    Timestamps are remapped from your clip's length to a 4–15 second target, so the result is generatable as-is.

    Remapped to 4–15 s

One extraction a week without an account, three when signed in, then 5 credits each.

Questions people actually ask

MiniMax H3 Max video to prompt FAQ

What to know before you turn a clip back into a MiniMax H3 Max prompt

What is MiniMax H3 Max video to prompt?

Reverse engineering for prompts: it watches a reference clip, frames and audio both, and returns the structured three-field prompt that could produce a similar one, with timestamps refitted to H3's 4-to-15-second range.

Will I get the exact same video back?

No. H3 exposes no seed to replay, and a prompt is a description, not a recording. Expect the same kind of shot, not the same shot.

Why not use a generic video to prompt tool?

Two reasons. Most sample keyframes and never hear the clip, so the sound is invented or missing. And they return prose, while H3 wants three labelled fields in a fixed order.

Should I just use the clip as a reference instead?

If you want to keep its framing, lighting or soundtrack — yes. Reference mode preserves rather than approximates. This page is for learning and re-creating, not preserving.

My source is forty seconds long. What happens?

You pick a section on the timeline, or let the beats be compressed proportionally. H3 tops out at 15 seconds, and every timestamp is recomputed to the target you choose.

Does it transcribe the dialogue?

Yes — exact words, speaker IDs, a delivery note and a language tag, in the format H3 reads. Eleven dialogue languages have stable support.

What do I do with the prompt?

Copy any field on its own, or send the whole thing into the text to video generator with one click. It is plain text, so it also works in the MiniMax API, Hailuo or ComfyUI.

Is it legal to recreate someone else's video?

Style, generally yes; a specific work or a real person's likeness, no. The section above draws the line, and the full version is on our responsible use page.

What can I paste or upload, and how long can it be?

A direct public URL to an MP4 or MOV, or a file up to 50 MB and 60 seconds. Anything past 15 seconds gets the segment picker, since that is H3's ceiling.

Does it work on clips with no dialogue?

Yes. It fills the soundscape and score fields and leaves the dialogue out — silence in a field is information too.

Is it free?

There is a free allowance every week: one extraction with no account, three once you sign in. Past that it costs 5 credits, roughly a thirteenth of what generating a video costs, because reading a clip is far cheaper than making one. Rate limits apply to everybody either way.

Can it go the other way — idea to prompt?

That is the prompt generator, this tool's mirror image. Write there when you have an idea and no footage; extract here when you have footage and no words.

Paste the clip, read it back as a prompt.

One free MiniMax H3 Max extraction a week without an account, and what you copy is yours to keep.

Generate free · queue

Written and maintained by the MiniMax H3 Max AI Video Generator editorial teamPublished Last updated