An independent comparison of AI video models
There are too many video models, and no honest way to pick one.
Every platform hands you a list of twenty and calls it choice. None of them shows you the same shot rendered by all of them, with what it cost and how long you waited. That is the whole of this site.
9 inputs · 7 models · 59 takes · free, no account, nothing to sign up for
the input
H3 Max
H3 ref2v
Kling v3 Pro
Wan 2.7
Seedance 2.0The problem
You have a character sheet, a product photo, a script and a deadline. What you do not have is any basis for choosing between Seedance and H3 and Kling, because nothing you can read tells you which one does your job.
Leaderboards rank quality with an Elo, voted by a crowd on prompts that are not yours. That is a real measurement of something, and it is not the thing you need. It cannot tell you which model holds a character across a cut, which one actually uses the voice you recorded instead of faking a sound-alike, which one silently returns eight seconds when you asked for ten, or which one refuses your reference photo outright.
Catalogs are worse. They list what a model claims to accept. They do not tell you what happens when you send it.
What this is
The same establishing frame, the same prompt, the same requested length, sent to every model, with the outputs playing side by side on one scrubber. Then a fixed set of deterministic checks on each take, and the numbers shown beside the video rather than summarised into a score.
- Cost and wall-clock on every take. Not list price in a table somewhere — what this shot actually cost and how long it actually took.
- Refusals recorded as refusals. A model that rejects a photoreal human face is telling you something you need before you build on it.
- Input-contract truth. The duration caps, reference limits and silent resolution substitutions that no schema mentions.
- Two checks nobody else runs. Does a model give you the same result twice, and does it hold a character across separate calls.
The models
Every one of them run on every input, on the same day, through the same provider. 2 of those calls came back as a refusal, and those are on the site too. Total spent producing what you are about to look at: $49.44.
What this is not
Not an arena — there is no vote and no Elo. Not a benchmark suite — there is no composite score ordering these models from best to worst, because no such ordering survives contact with a real brief. Not a storefront — nothing here is affiliated, sponsored, or paid for by anyone whose model appears on it.
It is one person running the same inputs through every model and publishing everything, including the parts that make the method look bad. Where a measurement fails to separate the field, the site says so.