The models
Shape is what the model's schema accepts; price is the provider's list on the date shown; every other number comes from the takes on this site and nothing else. Quality rankings live on the leaderboards linked from each card โ this page does not repeat them.
MiniMax H3 Max, image-to-video
Post-trained by fal for prompt adherence. One image in, no multi-reference. Continues the frame rather than re-staging it, which is why it holds identity better than models with more reference slots.
MiniMax H3, reference-to-video
The multi-reference sibling, and the only model here that embeds a supplied recording rather than imitating it. Re-stages the scene from the references, so adherence is better and continuity with the frame is worse.
Kling v3 Pro, image-to-video
Highest resolution here and the most expensive ten-second option. Weights the start frame far above its element references.
Wan 2.7, reference-to-video
No audio parameter but returns audio anyway. Inserts its own cuts inside a single shot and mixes hot โ two things no schema tells you.
Seedance 2.0, reference-to-video
Works where 2.5 refused every attempt. Refuses photoreal human references outright as possible likenesses, including your own face. Accepts an audio reference and synthesises a sound-alike rather than using it.
Veo 3.1, reference-to-video
Cleanest short render and the fastest multi-reference model here. This endpoint renders exactly eight seconds: a five-second input comes back at eight, a ten-second block comes back at eight. A six-second request is a 422.
LTX-2 19B, image-to-video
The price floor, and it shows on anything with dialogue or fast motion. Its portrait_16_9 preset delivers 576ร1024, not the 768ร1344 the name implies.
H3 reference-to-video + recorded audio
The same endpoint with a cast recording passed as the audio reference. The only route here that puts the actual recording in the output rather than a sound-alike โ cross-correlation 0.94 and 0.96 against 0.1โ0.3 for everything else.
Seedance 2.0 + recorded audio
Takes the recording as a reference and says every line โ in a voice that is not the recording. The check that separates this from the H3 route is the one nobody else runs.
Seedance 2.5, omni-reference
Estimated price, not measured. Credit-priced, so this is an estimate rather than a measurement. The route bills 6.5 credits per second at 720p โ 32.5 for a 5 s take, 65 for 10 s, confirmed against the ledger. Higgsfield credits work out at roughly $0.033 to $0.05 each depending on plan and whether they are bought in a pack, which puts this model somewhere between $0.21 and $0.46 per second, most likely $0.25โ0.33. Three reasons it stays out of the cost column: subscription credits are bundled, so the marginal cost of one more take is arguably zero until the allowance runs out, while every other price here is metered per call; the published tiers do not cover every plan; and the credit packs are currently discounted 40-odd percent, so the rate moves. Worth noting the direction anyway โ fal lists the same model at $0.473/s at 720p, so this route is cheaper, not marked up.
The endpoint takes a mode rather than one URL per input shape, which is why 2.0-shaped parameters failed against it. More to the point: the seven-reference character bible that the metered route refused as a "likeness of real people" went through here first try, all four scripted lines heard. The refusal was the route, not the model.