Consistency
Two questions no leaderboard asks. Does a model give you the same result twice, and does it hold a character across separate calls? Both are measured here, both with plain thresholds you are welcome to argue with.
Repeatability
The same input, rendered 3 times on the same model, no seed forced. The number is the coefficient of variation of motion across replicates — spread over mean, so a busy model and a calm one compare honestly. Lower is steadier. Only models cheap and fast enough to run repeatedly are here; running five Veo takes to learn what two H3 Max takes tell you is not a good trade.
| Model | Motion CV | Inputs | Per input |
|---|---|---|---|
| H3 Max | 0.164 | 1 | Stylised animation 0.16 |
| LTX-2 | 0.773 | 1 | Stylised animation 0.77 |
Identity across calls
One character, two separate generations, nothing carried between them but the reference images. Each appearance is located with an open-vocabulary detector — a frame where the character cannot be found at all is rejected rather than scored, so a bad render cannot quietly count as a consistent one — then cropped, embedded, and compared to that character's own centroid across calls. Higher is steadier. Rejections are shown because a model whose output is unrecognisable is telling you something too.
| Model | Character | Mean similarity | Worst | Appearances | Detail |
|---|---|---|---|---|---|
| Kling v3 Pro | The baker | 0.975 | 0.975 | 2 / 2 frames kept | 2 calls compared to their own centroid · 10/10 sampled frames passed the fidelity gate |
| H3 ref2v | The baker | 0.972 | 0.972 | 2 / 2 frames kept | 2 calls compared to their own centroid · 10/10 sampled frames passed the fidelity gate |
| H3 ref2v + voice | The baker | 0.969 | 0.969 | 2 / 2 frames kept | 2 calls compared to their own centroid · 10/10 sampled frames passed the fidelity gate |
| H3 Max | The baker | 0.967 | 0.967 | 2 / 2 frames kept | 2 calls compared to their own centroid · 10/10 sampled frames passed the fidelity gate |
| Wan 2.7 | The baker | 0.965 | 0.965 | 2 / 4 frames kept | 2 calls compared to their own centroid · 8/10 sampled frames passed the fidelity gate |
| Veo 3.1 | The baker | 0.962 | 0.962 | 2 / 2 frames kept | 2 calls compared to their own centroid · 10/10 sampled frames passed the fidelity gate |
| Seedance + voice | The baker | 0.962 | 0.962 | 2 / 2 frames kept | 2 calls compared to their own centroid · 10/10 sampled frames passed the fidelity gate |
| LTX-2 | The baker | 0.961 | 0.961 | 2 / 2 frames kept | 2 calls compared to their own centroid · 10/10 sampled frames passed the fidelity gate |
| H3 Max | The paper fox | 0.955 | 0.955 | 2 / 2 frames kept | 2 calls compared to their own centroid · 10/10 sampled frames passed the fidelity gate |
| Seedance 2.0 | The baker | 0.947 | 0.947 | 2 / 2 frames kept | 2 calls compared to their own centroid · 10/10 sampled frames passed the fidelity gate |
| LTX-2 | The paper fox | 0.941 | 0.941 | 2 / 3 frames kept | 2 calls compared to their own centroid · 9/10 sampled frames passed the fidelity gate |
| Veo 3.1 | serial-flyer | 0.905 | 0.905 | 2 / 6 frames kept | 2 calls compared to their own centroid · 6/10 sampled frames passed the fidelity gate |
Read this one carefully. Every model scored here lands in a narrow band, and a metric that cannot separate the field is a finding about the metric before it is a verdict on the models. Embedding a cropped stylised character captures "brown round cartoon animal" far more strongly than it captures which brown round cartoon animal, so the numbers agree more than the clips do. Treat the ordering as weak evidence and the rejections as strong evidence.
Not scored, and why that matters
A row here means the character could not be found in enough frames to compare across calls. That has two very different causes and this check cannot tell them apart: the model may have drifted so far that the character is no longer recognisable, or the character may simply be small, dark or partly out of frame and the open-vocabulary detector weak on it. Both are worth knowing; neither is a quality score.
- H3 Max · serial-flyer — only one input produced a usable appearance, so there is nothing to compare across calls
- H3 ref2v · serial-flyer — the entity was not found in any sampled frame — no score possible
- Kling v3 Pro · serial-flyer — only one input produced a usable appearance, so there is nothing to compare across calls
- Wan 2.7 · serial-flyer — only one input produced a usable appearance, so there is nothing to compare across calls
- Seedance 2.0 · serial-flyer — only one input produced a usable appearance, so there is nothing to compare across calls
- LTX-2 · serial-flyer — the entity was not found in any sampled frame — no score possible
- H3 ref2v + voice · serial-flyer — only one input produced a usable appearance, so there is nothing to compare across calls
- Seedance + voice · serial-flyer — only one input produced a usable appearance, so there is nothing to compare across calls
- H3 ref2v · The paper fox — only one input produced a usable appearance, so there is nothing to compare across calls
- Kling v3 Pro · The paper fox — only one input produced a usable appearance, so there is nothing to compare across calls
- Wan 2.7 · The paper fox — only one input produced a usable appearance, so there is nothing to compare across calls
- Seedance 2.0 · The paper fox — only one input produced a usable appearance, so there is nothing to compare across calls
- Veo 3.1 · The paper fox — only one input produced a usable appearance, so there is nothing to compare across calls
The pairs
Inputs that share a character. Everything else about them differs.



