Model Bakeoffrendered on fal · 2026-09-03

Consistency

Two questions no leaderboard asks. Does a model give you the same result twice, and does it hold a character across separate calls? Both are measured here, both with plain thresholds you are welcome to argue with.

Repeatability

The same input, rendered 3 times on the same model, no seed forced. The number is the coefficient of variation of motion across replicates — spread over mean, so a busy model and a calm one compare honestly. Lower is steadier. Only models cheap and fast enough to run repeatedly are here; running five Veo takes to learn what two H3 Max takes tell you is not a good trade.

ModelMotion CVInputsPer input
H3 Max0.1641 Stylised animation 0.16
LTX-20.7731 Stylised animation 0.77

Identity across calls

One character, two separate generations, nothing carried between them but the reference images. Each appearance is located with an open-vocabulary detector — a frame where the character cannot be found at all is rejected rather than scored, so a bad render cannot quietly count as a consistent one — then cropped, embedded, and compared to that character's own centroid across calls. Higher is steadier. Rejections are shown because a model whose output is unrecognisable is telling you something too.

ModelCharacterMean similarityWorstAppearancesDetail
Kling v3 ProThe baker 0.9750.975 2 / 2 frames kept2 calls compared to their own centroid · 10/10 sampled frames passed the fidelity gate
H3 ref2vThe baker 0.9720.972 2 / 2 frames kept2 calls compared to their own centroid · 10/10 sampled frames passed the fidelity gate
H3 ref2v + voiceThe baker 0.9690.969 2 / 2 frames kept2 calls compared to their own centroid · 10/10 sampled frames passed the fidelity gate
H3 MaxThe baker 0.9670.967 2 / 2 frames kept2 calls compared to their own centroid · 10/10 sampled frames passed the fidelity gate
Wan 2.7The baker 0.9650.965 2 / 4 frames kept2 calls compared to their own centroid · 8/10 sampled frames passed the fidelity gate
Veo 3.1The baker 0.9620.962 2 / 2 frames kept2 calls compared to their own centroid · 10/10 sampled frames passed the fidelity gate
Seedance + voiceThe baker 0.9620.962 2 / 2 frames kept2 calls compared to their own centroid · 10/10 sampled frames passed the fidelity gate
LTX-2The baker 0.9610.961 2 / 2 frames kept2 calls compared to their own centroid · 10/10 sampled frames passed the fidelity gate
H3 MaxThe paper fox 0.9550.955 2 / 2 frames kept2 calls compared to their own centroid · 10/10 sampled frames passed the fidelity gate
Seedance 2.0The baker 0.9470.947 2 / 2 frames kept2 calls compared to their own centroid · 10/10 sampled frames passed the fidelity gate
LTX-2The paper fox 0.9410.941 2 / 3 frames kept2 calls compared to their own centroid · 9/10 sampled frames passed the fidelity gate
Veo 3.1serial-flyer 0.9050.905 2 / 6 frames kept2 calls compared to their own centroid · 6/10 sampled frames passed the fidelity gate

Read this one carefully. Every model scored here lands in a narrow band, and a metric that cannot separate the field is a finding about the metric before it is a verdict on the models. Embedding a cropped stylised character captures "brown round cartoon animal" far more strongly than it captures which brown round cartoon animal, so the numbers agree more than the clips do. Treat the ordering as weak evidence and the rejections as strong evidence.

Not scored, and why that matters

A row here means the character could not be found in enough frames to compare across calls. That has two very different causes and this check cannot tell them apart: the model may have drifted so far that the character is no longer recognisable, or the character may simply be small, dark or partly out of frame and the open-vocabulary detector weak on it. Both are worth knowing; neither is a quality score.

The pairs

Inputs that share a character. Everything else about them differs.