fal MiniMax H3 Max: the Rate Card, and What It Actually Makes

Both post-trained builds, all five endpoints and every published rate — the table fal spreads across five pages. Then the two things a rate card cannot tell you: what it makes, and what the rate does not cover.
Rates read from fal's own model pages on 4 September 2026. Every figure below quadruples on 7 September.

H3 Max Turbo

minimax/h3-max-turboCheaper · newer
text-to-video
A prompt, up to 7,000 characters. 5–15 whole seconds at 24 fps, sound generated in the same pass.
480p$0.025/s
768p$0.04/s
image-to-video
A prompt and a still to start from.
480p$0.025/s
768p$0.04/s
reference-to-video
Not published for this model. Guess the URL from the pattern and you get a 404 — worth knowing before you write the client.
480p
768p

fal describes it as a post-trained variant of MiniMax H3 tuned for stronger prompt adherence and better aesthetics, co-optimised with its inference stack for higher throughput. It is the default engine on this site, which is why the numbers further down are measured on it.

MiniMax H3 Max

minimax/h3-maxThe launch build
text-to-video
A prompt, up to 7,000 characters. Same durations, same frame shapes, same in-pass audio.
480p$0.05/s
768p$0.08/s
image-to-video
A prompt and a first frame, a last frame, or both.
480p$0.05/s
768p$0.08/s
reference-to-video
Up to 12 files total: reference images, clips of 2–15 s (≤3, ≤15 s combined) and audio on the same limits. Never audio alone. Cited in the prompt as Image 1, Video 1, Audio 1.
480p$0.05/s + tokens
768p$0.08/s + tokens

The build fal and MiniMax released together on 26 August 2026 — twice Turbo's rate at both resolutions, and the only one of the two with a reference task. That reference input is billed on its own meter and it can cost more than the clip; the arithmetic is further down.

Those are fal's list rates, so this table still reads correctly next month. A launch discount is running at a quarter of them until 7 September 2026 — both figures are in the calculator below.
This site is an independent third-party interface. It is not affiliated with fal or with MiniMax, we earn nothing from sending you there, and the figures above are quoted from fal's public pages, which can change without us noticing: fal's MiniMax H3 Max page.

Four out of 299, ours

What Does fal MiniMax H3 Max Actually Make?

We put 300 prompts through minimax/h3-max-turbo/text-to-video on 4 September 2026 — five seconds, 480p, 16:9, sound generated in the same pass — and 299 came back. These four are straight off that run: no grade, no upscale, no second pass, and the sound is the model's.

Twins in matching tracksuits demonstrate a briefcase in a rain-lashed car park at night, shot like a 1980s television advert00:00:05:00
A 1980s advert for a product that does not exist. One prompt, one request.
A young news anchor at a desk reads a bulletin while two people work in the newsroom behind, a lower third across the bottom of frame00:00:05:00
Rolling news, lower third and all. The voice came out with the picture.
A pale fish with too many fins grooming itself at the mouth of a limestone cave under flat overcast light00:00:05:00
Natural history for an animal that has never existed. Rain, no music.
A paper flower opening slowly against a deep blue gradient as a channel name resolves out of settling dust00:00:05:00
A station ident, down to the scanlines and the library music loop.

299 / 300

Clips that landed

5.03 s

Median, submit to file

124

Frames on every one

24 min 55 s

Of finished programme

Every number below the fold comes off this run rather than a launch post — the speed distribution, the frame counts, the file sizes and what it cost. The other 295 are playing on the channel page, each with the prompt that made it.

The short answer

What Is fal's MiniMax H3 Max?

A model MiniMax and fal post-trained together, hosted on fal. MiniMax published the open-weights base — MiniMax H3 — and fal post-trained it for generation speed, releasing the result on 26 August 2026. It is the model behind the endless-AI-channel experiments people have been posting all week, and the reason those are possible at all is on this page: it finishes a clip faster than the clip plays.
It renders 480p or 768p, never 2K, in whole seconds from 5 to 15 at 24 fps, with sound generated in the same pass as the picture rather than dubbed on afterwards. Neither party has published the post-train's weights; the base model's are public.

There are two of them now, and the cheaper one is newer. Read on 4 September 2026, fal lists minimax/h3-max-turbo at $0.025/s at 480p and $0.04/s at 768p, against $0.05 and $0.08 for minimax/h3-max — an exact 2× gap at both resolutions. Turbo publishes no reference endpoint.
If you only take one thing from the naming: the word Turbo means three different things in this ecosystem, and two of them are for sale at a 2× price gap.

You can use it right now without any of the above. No fal account, no API key, no payment method, no code — text to video and image to video are the same two tasks as the table, in a browser, and the first clip does not ask for a card.
That is a genuinely different question from which per-second rate is lower, and the next section is the arithmetic for deciding it rather than an argument about it.

Post-trains on fal
2

Post-trains on fal

Endpoints between them
5

Endpoints between them

Rate rise on 7 Sep

Rate rise on 7 Sep

Our median 5-second clip
5.03 s

Our median 5-second clip

  • No 2K on either build — 768p is the ceiling
  • No four-second clips; five is the floor
  • No video-to-video endpoint on fal for either build
  • No published weights for either post-train

The rest of the bill

What Does fal MiniMax H3 Max Cost to Run Yourself?

A per-second rate prices one render, and nobody buys one render — you buy a clip you are willing to keep, which takes however many attempts it takes. Set the dials to your real job. The figure is fal’s, at fal’s published rates, and what sits under it is the part of the job the rate does not buy.

Model

Resolution

10
5s
3×

Your estimate, not our measurement. We have no honest figure for how many attempts a person needs — so this dial is yours. Set it to 1 if your prompt is settled.

What fal bills

30 renders × 5 s = 150 billed seconds

List rate · from 7 Sep 2026

$3.75

$0.38 a finished clip

Launch discount · until 7 Sep 2026

$0.94

$0.09 a finished clip

Same job. Four times the money. Three days apart — and the higher figure is the one that is true for the rest of the year.

What that figure does not include

  • A fal account, a card on file and an API key. A per-second rate is not a price a person can pay — it is a price a program can pay.
  • The client that submits, polls the queue, retries the failures and downloads the files. Ours ran to 24,000 characters, and it still stopped at clip 300 on a 403.
  • Somewhere to put them. 299 clips is 735 MB at 480p, before anyone has watched one.
  • The time before the first render. All of the above happens before you see a frame. Here it is one field and about five seconds.

At this size, the rate is not what you are paying.

A job this size is a couple of dollars of compute either way. What it actually costs you is the account, the key, the client code and the afternoon they take — spent to save less than the price of lunch.
The same two tasks run here in a browser: one field, five seconds, sound in the same pass, and the first clip does not ask for a card. If you would rather have the API anyway, it will still be there afterwards.

Make the first one free

This prices fal's API only, at fal's published rates. Here is ours for the same render, so the one number you came to compare against is not the one number this page withholds. A five-second H3 Max Turbo clip costs $0.43 at 480p and $0.66 at 768p here on Pro, against $0.125 and $0.20 at fal's list rate — and a quarter of those until 7 September.
So we are about three and a half times the compute, and about thirteen times it this week. That difference is the browser, the prompt tooling, the engine picker, the storage, the automatic refund when a render errors, and not needing an account to begin — the whole rate card is here. On 7 September fal's figures step up to the list row and ours do not move, so the same comparison reads 3× instead of 13× with neither side touching a price.

7 September 2026

What Changes for fal MiniMax H3 Max on 7 September

fal states the end of its launch discount on its own model pages: the discount ends September 7, after which 480p is $0.05 per second and 768p is $0.08. Read 4 September 2026. One sentence, both builds, every rate.

If you are costing a project on the price you were quoted this week, you are costing it wrong. The rate card has an expiry date on it, and the number on the far side is the one you will live with.

  • H3 Max Turbo · 480p

    Until 7 September
    $0.00625/s
    From 7 September
    $0.025/s
    Change
  • H3 Max Turbo · 768p

    Until 7 September
    $0.01/s
    From 7 September
    $0.04/s
    Change
  • MiniMax H3 Max · 480p

    Until 7 September
    $0.0125/s
    From 7 September
    $0.05/s
    Change
  • MiniMax H3 Max · 768p

    Until 7 September
    $0.02/s
    From 7 September
    $0.08/s
    Change
  • One 5 s clip · Max at 768p · first try

    Until 7 September
    $0.10Worked example
    From 7 September
    $0.40
    Change
  • 300 clips · 25 min · Turbo at 480p

    Until 7 September
    $9.38We ran it: $9.28
    From 7 September
    $37.50
    Change
    The run

Where this leaves you

  • A quote you gave a client this week is a quarter of what the work will bill next week
  • A pipeline costed at the launch rate has a 4× step change in it, on a date, not a negotiation
  • The discount applies to renders you discard as well as the ones you ship
  • Reference-to-video moves too, and its token charge sits on top of whichever rate is live

None of this is a criticism of fal — a launch discount is a launch discount, and they published the end date up front. It is only a problem if you budget as though it were the price.

What the same three days do to our price

Nothing — and that is the only part of this page we control, so it is worth being precise about what it means for the comparison you are running:

  • The comparison improves fourfold on Sunday, and we do nothing

    A five-second 768P clip is $0.05 of fal compute this week against $0.66 here. From 7 September it is $0.20 against the same $0.66. Nobody changes a price and the gap goes from about 13× to about 3×, because this rate card was built on fal's list price rather than on its promotion — the derivation is on the pricing page.

  • Which makes this the worst week of the year to read this page

    We are publishing the 13× version of that comparison during the four days it is true, rather than waiting until Monday when it reads 3×. A page that only quotes itself in the flattering week is a page you cannot plan against.

  • And you do not have to take any of it on trust

    The first clip does not ask for a card, so none of this pricing has to be your problem before you know whether the output is worth paying for at all.

Sources and method

299 clips, our files, our ffprobe

How Fast Is fal MiniMax H3 Max, Measured?

Every page about this model quotes the same launch figure. We generated a 300-clip programme tape on minimax/h3-max-turbo/text-to-video on 4 September 2026 — 5 seconds, 480p, 16:9, one request each — and measured what came back. 299 landed. This is the distribution, which is the part nobody publishes.

The clips arrive faster than they play. A median of 5.03 seconds of wall clock for 5.167 seconds of video. That is the whole reason the endless-channel experiments work, reproduced on ordinary queued requests from an ordinary account — not on a demo rig.

Submit → file on disk
p50 5.03 s · p95 15.42 s · fastest 3.08 s · slowest 43.50 s
fal's own inference time
p50 0.452 s, self-reported in the response · everything else is queue, encode and download
Frames returned
124 at 24 fps on all 299 — 5.167 s of video from a 5-second request
480p at 16:9
832×480 on all 299 — an aspect of 1.733, not 1.778
Audio
Present on all 299, generated in the same pass as the picture
File size
2.46 MB average at 5 s / 480p
What it cost us
$9.28 for 299 clips at the launch rate — the same tape is $37.50 from 7 September

Plan a pipeline around the p95, not the median. One clip in twenty took three times the median and one took 43 seconds, so a queue sized on "about five seconds" stalls. That is the kind of thing you only learn by running a few hundred, and it is the reason this section exists rather than a quote from a launch post.
You ask for five seconds and you are given 5.167. Frame counts land on a 17n+5 grid, so a whole-second request is rounded to the nearest grid point — 124 frames on every one of the 299, no exceptions. A small thing until you are cutting to music or billing by the second.
For contrast, the launch build was probed separately the day before on the same frame shapes: 6.6 s for a 5-second 768p clip, 21.2 s for a 15-second one, 4.4 s at 480p. Three requests rather than 299 — spot checks, not a distribution, and labelled as such.
The 299 clips were not thrown away. They are a programme tape, and they are playing on our channel page with the prompt that made each one printed underneath it.

The bill that surprises people

Why Reference to Video Costs More Than the Clip It Makes

Only minimax/h3-max publishes this task, and it is the one endpoint where the per-second rate is not the bill. Reference input is metered separately as tokens: the first 4,096 free, then $0.02 per 1,000. If you are pricing character consistency into a quote, this is the line you are about to miss.

A reference clip's token count follows the resolution you render at, not the resolution you uploaded. The identical 5-second reference costs 3.1× more against a 768p render than against a 480p one — so raising output quality silently raises your input bill too.

First 4,096 tokens
Free, on every request
After that
$0.02 per 1,000 tokens
A 1024×1024 image
1,024 tokens — so roughly four reference images ride free
Reference video, 480p render
2,886 tokens a second — $0.21 for a 5-second reference, $0.78 for 15
Reference video, 768p render
7,459 tokens a second — $0.66 for a 5-second reference, $2.16 for 15
Reference audio
About 80 tokens a second — $0.0016 for 5 seconds, $0.024 for 15
What it accepts
≤12 files total · clips 2–15 s, ≤3, ≤15 s combined · audio on the same limits · never audio alone

Worked through, because the shape of it catches people: a 5-second 768p clip with one 5-second reference video is $0.40 of output plus $0.66 of reference input. The reference costs more than the video it is steering. At 15 seconds it is $1.20 plus $2.16. Now put that through the retry dial above — the reference is re-billed on every attempt, not just the keeper.
Reference-to-video is not available on the MiniMax H3 Max first-party API yet; fal exposes a separate reference endpoint. On this site, reference-to-video runs on MiniMax H3 — the base model — at one posted rate with no separate input meter, and the engine badge on that page says which model produced every render.

Questions people actually ask

fal MiniMax H3 Max FAQ

What it is, what it lists at, when the API is the right call, and what the reference endpoint really bills

What is fal MiniMax H3 Max?

MiniMax H3 Max is a post-trained version of MiniMax's open-weights MiniMax H3 model, tuned for generation speed and released on fal on 26 August 2026. It renders 480p or 768p — there is no 2K — in whole seconds from 5 to 15 at 24 frames a second, with sound generated in the same pass as the picture rather than dubbed on afterwards. The weights of the post-train have not been published by MiniMax or by fal; the base model's weights have, at MiniMaxAI/MiniMax-H3 on Hugging Face.

How much does MiniMax H3 Max cost on fal?

Read on 4 September 2026, fal's list rates are $0.05 per second of output at 480p and $0.08 at 768p for minimax/h3-max, and $0.025 and $0.04 for minimax/h3-max-turbo. A launch discount is running at a quarter of those figures until 7 September 2026, after which fal's own model pages state that 480p returns to $0.05 per second and 768p to $0.08. Two things that rate does not include: every render you discard is billed the same as the one you keep, and reference-to-video adds a separate token charge for the reference input on top of the per-second rate.

Is it cheaper to use the fal API or a hosted MiniMax H3 Max interface?

Per second of output, the API is cheaper, and for a large scripted batch with a settled prompt it is the right tool — we would tell you to use it. The comparison changes for smaller jobs, because the API bills every attempt at the same rate as the keeper and because the price you are quoted today is a launch discount that ends on 7 September 2026. A ten-clip job at five seconds is a few dollars of compute either way, and at that size the decision is not the money — it is whether you want to hold a fal account, an API key, a payment method and code that submits, polls, retries and downloads, or whether you would rather have the clip in the next five seconds. We wrote that code to generate 300 clips, and it still stopped at clip 300 on an exhausted balance.

Is minimaxh3max.video cheaper than the fal API?

No, and the figures are better coming from us than from your invoice. A five-second H3 Max Turbo clip is $0.43 at 480p and $0.66 at 768p here on the Pro plan, against $0.125 and $0.20 for the same renders at fal's list rate — about three and a half times — or about thirteen times while the launch discount runs to 7 September 2026. fal sells raw inference and prices it accordingly. What the difference buys is everything between a rate and a finished clip: no fal account, no card, no API key and no client code that submits, polls, retries and downloads; the prompt tooling that decides how many renders you pay for in the first place; the engine picker, the storage and the moderation; and a render that errors returns its credits automatically rather than being argued about. If you are running a scripted batch with a settled prompt, use the API — the calculator on this page says the same thing once you slide it past a few hundred renders. Below that, what decides your bill is how many attempts a clip takes, not what one render lists at.

What is the difference between fal's H3 Max Turbo and MiniMax H3 Max?

They are two separate models on two separate fal model pages, at an exact 2x price gap at both resolutions. fal describes minimax/h3-max-turbo as a post-trained variant of MiniMax H3 tuned for stronger prompt adherence and better aesthetics, co-optimised with its inference stack for higher throughput. The one thing the more expensive minimax/h3-max still has that Turbo does not is a reference-to-video endpoint: Turbo publishes text-to-video and image-to-video only, and a request to a reference path under it returns 404.

Does fal's MiniMax H3 Max support reference to video?

minimax/h3-max does; minimax/h3-max-turbo does not. The endpoint accepts up to 12 files in total — reference images, clips of 2 to 15 seconds each up to 3 of them and 15 seconds combined, and audio on the same limits — and it will not accept audio alone. Billing is unusual: reference input is charged as tokens on top of the per-second output rate, with the first 4,096 tokens free and $0.02 per 1,000 after that. A reference clip's token count follows the resolution you render at, not the resolution you uploaded, so the same 5-second reference is $0.21 against a 480p render and $0.66 against a 768p one — which is more than the five seconds of video it is steering.

How fast is MiniMax H3 Max on fal?

We measured 299 clips on minimax/h3-max-turbo/text-to-video on 4 September 2026, each one 5 seconds at 480p and 16:9. From submitting the request to the finished file on disk, the median was 5.03 seconds and the 95th percentile 15.42 seconds, with the fastest at 3.08 and the slowest at 43.50. fal's own reported inference time had a median of 0.452 seconds; the rest is queue, encode and download. Every one of the 299 returned 124 frames at 24 frames a second, which is 5.167 seconds of video — so the median clip arrived faster than it plays. The tail matters as much as the median: one clip in twenty took three times the median, which is the number to plan a pipeline around.

Can I use MiniMax H3 Max without a fal API key?

Yes. minimaxh3max.video is an independent third-party interface that runs both post-trained builds, so text-to-video and image-to-video work here from a browser with no API key, no payment method and no code. This site is not affiliated with fal or with MiniMax, and the fal rates quoted on this page are fal's own published figures for its platform, not ours.

Is MiniMax H3 Max free on fal?

fal has run a free daily allowance on its model pages, but the terms have moved since launch and we have not re-verified them, so the only accurate answer is to read whatever fal's own page shows at the time you open it. What we can state is what happens here: a free account gets real MiniMax H3 clips without a card, and the current terms are set out on our free page.

Render One on H3 Max Before You Price It

Both post-trained builds run here in a browser — no key, no card for the first clip, and the engine picker tells you which one rendered what. Five to fifteen seconds, up to 768P, sound in the same pass as the picture. About five seconds from writing the line to watching the clip, which is the only part of this page you cannot get from a rate card.

Generate free · queue

Written and maintained by the MiniMax H3 Max AI Video Generator editorial teamPublished Last updated