Model Bakeoffrendered on fal ยท 2026-09-03

The models

Shape is what the model's schema accepts; price is the provider's list on the date shown; every other number comes from the takes on this site and nothing else. Quality rankings live on the leaderboards linked from each card โ€” this page does not repeat them.

MiniMax H3 Max, image-to-video

minimax/h3-max/image-to-video
1 framenative audio5โ€“15 s768p native
price$0.08/s
wait p50 / max7s / 13s
checks all green5 of 9
lines heard100%
refused0
spent here$5.20
repeatability0.16 cv
identity0.96

Post-trained by fal for prompt adherence. One image in, no multi-reference. Continues the frame rather than re-staging it, which is why it holds identity better than models with more reference slots.

MiniMax H3, reference-to-video

minimax/h3/reference-to-video
โ‰ค9 imagesโ‰ค3 video refsโ‰ค3 audio refs5โ€“15 s768p native
price$0.06/s
wait p50 / max3:34 / 10:24
checks all green3 of 8
lines heard100%
refused0
spent here$3.32
identity0.97

The multi-reference sibling, and the only model here that embeds a supplied recording rather than imitating it. Re-stages the scene from the references, so adherence is better and continuity with the frame is worse.

Kling v3 Pro, image-to-video

fal-ai/kling-video/v3/pro/image-to-video
1 framecharacter elementsnative audio1โ€“15 s1080p
price$0.196/s
wait p50 / max3:34 / 6:39
checks all green6 of 8
lines heard90%
refused0
spent here$9.80
identity0.97

Highest resolution here and the most expensive ten-second option. Weights the start frame far above its element references.

Wan 2.7, reference-to-video

fal-ai/wan/v2.7/reference-to-video
โ‰ค5 reference mediaaudio, no switch2โ€“10 s720p / 1080p
price$0.10/s
wait p50 / max2:32 / 5:27
checks all green4 of 8
lines heard100%
refused0
spent here$5.00
identity0.96

No audio parameter but returns audio anyway. Inserts its own cuts inside a single shot and mixes hot โ€” two things no schema tells you.

Seedance 2.0, reference-to-video

bytedance/seedance-2.0/reference-to-video
โ‰ค9 imagesโ‰ค3 audio refs4โ€“15 s720p
price$0.18/s
wait p50 / max3:38 / 3:56
checks all green4 of 6
lines heard100%
refused2
spent here$7.20
identity0.95

Works where 2.5 refused every attempt. Refuses photoreal human references outright as possible likenesses, including your own face. Accepts an audio reference and synthesises a sound-alike rather than using it.

Veo 3.1, reference-to-video

fal-ai/veo3.1/reference-to-video
โ‰ค3 imagesnative audio8 s only720p โ€“ 4K
price$0.40/s
wait p50 / max35s / 46s
checks all green0 of 5
lines heard100%
refused0
spent here$16.00
identity0.93

Cleanest short render and the fastest multi-reference model here. This endpoint renders exactly eight seconds: a five-second input comes back at eight, a ten-second block comes back at eight. A six-second request is a 422.

LTX-2 19B, image-to-video

fal-ai/ltx-2-19b/image-to-video
1 framenative audioopen weights576ร—1024 portrait
price$0.0018
wait p50 / max47s / 1:46
checks all green6 of 9
lines heard50%
refused0
spent here$2.92
repeatability0.77 cv
identity0.95

The price floor, and it shows on anything with dialogue or fast motion. Its portrait_16_9 preset delivers 576ร—1024, not the 768ร—1344 the name implies.

H3 reference-to-video + recorded audio

minimax/h3/reference-to-video
โ‰ค9 imagesrecorded takes as audio ref
price$0.06/s
wait p50 / max16:42 / 16:42
checks all green0 of 2
lines heard100%
refused0
spent here$1.52
identity0.97

The same endpoint with a cast recording passed as the audio reference. The only route here that puts the actual recording in the output rather than a sound-alike โ€” cross-correlation 0.94 and 0.96 against 0.1โ€“0.3 for everything else.

Seedance 2.0 + recorded audio

bytedance/seedance-2.0/reference-to-video
โ‰ค9 imagesrecorded takes as audio ref
price$0.18/s
wait p50 / max4:54 / 4:54
checks all green0 of 2
lines heard100%
refused0
spent here$3.60
identity0.96

Takes the recording as a reference and says every line โ€” in a voice that is not the recording. The check that separates this from the H3 route is the one nobody else runs.

Seedance 2.5, omni-reference

seedance_2_5 (mode: omni_reference)
โ‰ค7 images testedvideo + audio refs4โ€“30 s480pโ€“1080pmodes: t2v / omni / edit / extend
priceest. $0.25โ€“0.33/s
wait p50 / maxโ€” / 0s
checks all green1 of 2
lines heard75%
refused0
spent here$0.00

Estimated price, not measured. Credit-priced, so this is an estimate rather than a measurement. The route bills 6.5 credits per second at 720p โ€” 32.5 for a 5 s take, 65 for 10 s, confirmed against the ledger. Higgsfield credits work out at roughly $0.033 to $0.05 each depending on plan and whether they are bought in a pack, which puts this model somewhere between $0.21 and $0.46 per second, most likely $0.25โ€“0.33. Three reasons it stays out of the cost column: subscription credits are bundled, so the marginal cost of one more take is arguably zero until the allowance runs out, while every other price here is metered per call; the published tiers do not cover every plan; and the credit packs are currently discounted 40-odd percent, so the rate moves. Worth noting the direction anyway โ€” fal lists the same model at $0.473/s at 720p, so this route is cheaper, not marked up.

The endpoint takes a mode rather than one URL per input shape, which is why 2.0-shaped parameters failed against it. More to the point: the seven-reference character bible that the metered route refused as a "likeness of real people" went through here first try, all four scripted lines heard. The refusal was the route, not the model.