P—04 / ONE LINE TO CAMERA

Dialogue prompt for MiniMax H3 Max

One spoken sentence, written inside quotation marks, and the lips land on it.

The whole reason to use a model that generates audio with the picture, in 376 characters. It names who speaks, gives the exact words, says what she does after the line, and then describes the room and the microphone. Nothing about lip sync is requested anywhere in it — that comes from writing the line rather than describing it.

Steps
One shot
Engine
MiniMax H3 Max
Mode
Text to Video
Duration
5s
Ratio
16:9
Resolution
768P
Audio
Dialogue + kitchen foley
Cost
$0.40

Reference render — not generated on this site. Source: fal.ai

Output reference

A chef in a flour-dusted apron looks into the camera and speaks, a bright tiled kitchen behind herVideo
The clip
376 / 7000
  • Free clipThis run spends credits — the free clip is 5s · 480P · 16:9.

1 FREE CLIP · NO AUDIO

01 — Inside

One prompt, one workflow

The clip beside the verbatim source. Copy it into the console above, change what you need, generate.

Reference output

A chef in a flour-dusted apron looks into the camera and speaks, a bright tiled kitchen behind herVideo
The clip
The promptthe clipMiniMax H3 Max · Text to video
376 chars

The line

Speaker, the exact words in quotes, what happens after, then the light, the lens and the sound.

A chef in her forties looks straight into the camera in a bright open kitchen and says, "The secret is you never crowd the pan." She smiles and turns back to the stove. Clean daylight from a window, white tile, flour on her apron, a strand of hair loose. Medium close-up, 50mm, slight handheld. Sound: her voice close and clear, a gentle sizzle, a knife on a board off screen.

4 levers

Make it yours

What is safe to change. Most libraries publish only this list, which is why so many copied prompts come back worse than the original.

  1. 01

    The line itself

    Any sentence of roughly the same length works. Around 2.5 words a second is the natural pace, so eight to twelve words fit comfortably in five seconds. Write it as the words you want spoken, not as a description of them.

  2. 02

    The speaker

    "A chef in her forties" is age, gender and role in five words, which is enough to fix a delivery. Swap the role and the room together — a mechanic in a workshop, a florist at a bench — so the foley still matches.

  3. 03

    The beat after the line

    "She smiles and turns back to the stove" gives the clip somewhere to go once the sentence ends. Without it, five seconds contains three seconds of speech and two of a person looking at the camera.

  4. 04

    The off-screen sound

    "A knife on a board off screen" adds a second room without adding a second shot. It is the cheapest way to make a single-location clip feel like part of a larger place.

Three ways to break it

  • Describing the line instead of quoting it

    Write "she says something encouraging about cooking" and the model writes the sentence for you — rarely the one you wanted, and never the same one twice. The quotation marks are what turn dialogue into a fixed input.

  • Writing a tagged language in the wrong language

    If you mark a line as Japanese and then type it in English, the delivery comes back wrong. Whatever language you name, write the line in that language.

  • Adding a second speaker without splitting the time

    Five seconds holds one line at a natural pace. Two speakers need either a longer duration or two shorter lines, and asking for both inside five seconds compresses the delivery rather than extending the clip.

What it does

What the one line to camera template does

Every silent video model needs a second pass for this: generate the picture, then generate or record the voice, then sync them. This prompt is 376 characters and one pass, and the reason it works is almost entirely down to one piece of punctuation.

The quotation marks around "The secret is you never crowd the pan" turn the line from something the model has to invent into something it has to say. That distinction is the single most consequential thing in the whole prompt. Ask for "an encouraging line about cooking" and you get a different sentence every run, which means you cannot iterate: change the lighting and the words change too.

The length of the line is not accidental either. Eight words at a natural speaking pace runs a little over three seconds, which leaves room inside a five-second clip for her to arrive, say it and turn away. Dialogue that overruns is not truncated — the delivery is compressed to fit, and compressed delivery is where lip sync visibly fails first.

What comes after the line matters as much as the line. "She smiles and turns back to the stove" is the clip’s exit. A prompt that ends the moment the sentence ends leaves the model holding two seconds it has to fill, and what it fills them with is usually a face waiting.

The room is described in four short fragments: clean daylight from a window, white tile, flour on her apron, a strand of hair loose. Not one of them is a lighting instruction in the technical sense, and together they do the job of several. Daylight from a window sets direction and quality; white tile sets bounce; flour and loose hair are the two details that make the frame look inhabited rather than staged.

The sound line is written from the microphone’s point of view rather than the scene’s: her voice close and clear, a gentle sizzle, a knife off screen. Three sources at three distances. Because H3 Max writes the audio in the same pass as the picture, distance described is distance you hear — and it is what stops a talking-head clip from sounding like a voiceover laid over footage.

Everything here runs in text to video, on MiniMax H3 Max.

3 questions

One line to camera — common questions

  • 01

    How many words fit in a clip?

    About 2.5 a second at a natural pace, so roughly twelve in five seconds and thirty in fifteen. Past that the delivery is compressed rather than the clip extended.

  • 02

    Can I get two people talking?

    Yes, with more seconds and shorter lines. Tag each speaker and write each line in the language you tagged it with.

  • 03

    Do I need a reference image for a consistent face?

    For one clip, no — the description fixes it well enough. For the same face across several clips you need reference input, which runs on MiniMax H3 rather than H3 Max.

Copy it, change three things, run it.

Every character is on this page. 5s at 768P costs $0.40.

Generate free · queue

Written and maintained by the MiniMax H3 Max AI Video Generator editorial teamPublished Last updated