Model Bakeoffrendered on fal · 2026-09-03

Models › H3 Max

MiniMax H3 Max, image-to-video

minimax/h3-max/image-to-video

1 framenative audio5–15 s768p native
price$0.08/s
wait p50 / max7s / 13s
checks all green5 of 9
lines heard100%
refused0
spent here$5.20

Post-trained by fal for prompt adherence. One image in, no multi-reference. Continues the frame rather than re-staging it, which is why it holds identity better than models with more reference slots.

Everything it rendered

9 inputs, 11 takes. Click any one for the full request and its measurements.

Character serial, block 1 $0.80 · 13s lines 4/4 · synth voice · motion 3.35
Character serial, block 2 $0.80 · 13s lines 3/3 · synth voice · motion 3.61
One face, one line $0.40 · 7s lines 1/1 · motion 1.67
Two people, two lines $0.40 · 6s lines 2/2 · motion 2.36
Legible text on a sign $0.40 · 6s motion 3.44
Product on a turntable $0.40 · 6s motion 0.86
Fast camera, fast subject $0.40 · 6s motion 12.19 · clips
Stylised animation $0.40 · 7s · 3 replicates motion 0.64
The same fox, a different scene $0.40 · 8s motion 0.75

Repeatability

The same input sent again, more than once. The number is the spread of motion across replicates over its mean — lower is steadier.

InputReplicatesMotion CVDetail
Stylised animation30.1643 replicates · motion 0.64 / 0.89 / 0.76

Identity across calls

CharacterMean similarityAppearancesDetail
The baker0.96722 calls compared to their own centroid · 10/10 sampled frames passed the fidelity gate
serial-flyer1only one input produced a usable appearance, so there is nothing to compare across calls
The paper fox0.95522 calls compared to their own centroid · 10/10 sampled frames passed the fidelity gate

How every model compared →