One face, one line
A single human face has to hold its identity while the mouth matches a line the model voices itself.
2 models received different prompt text — show each
Each model is given the prompt its input contract calls for: a multi-reference model is told which reference is which character, a segmented model gets shot segments. The frame, the references and the requested length are identical throughout. This is the honest version of "the same prompt".
H3 Max, H3 ref2v, Wan 2.7, Veo 3.1, LTX-2
The woman leans forward slightly, looks into the camera and says: "I told you it would work. I just didn't say when." Then she picks up the mug and takes a sip. The camera holds steady. Natural morning light, room tone, no music.
Kling v3 Pro
[{"prompt":"The woman leans forward slightly, looks into the camera and says: \"I told you it would work. I just didn't say when.\" Then she picks up the mug and takes a sip. The camera holds steady. Natural morning light, room tone, no music.","duration":"5"}]- “I told you it would work. I just didn't say when.”
The images or videos provided may contain likenesses of real people or other private information that cannot be processed."
| H3 Max | H3 ref2v | Kling v3 Pro | Wan 2.7 | Seedance 2.0 | Veo 3.1 | LTX-2 | |
|---|---|---|---|---|---|---|---|
| Lines heard | 1/1 | 1/1 | 1/1 | 1/1 | — | 1/1 | 1/1 |
| Voice | model's own | model's own | model's own | model's own | — | model's own | model's own |
| Motion | 1.67 | 1.79 | 2.51 | 1.80 | — | 2.15 | 2.13 |
| Scene changes | 0 | 0 | 0 | 0 | — | 0 | 0 |
| Loudness / peak | -20.5 / -9.4 | -13.5 / -1.9 | -20.4 / -8 | -12.5 / -0.6 | — | -19.4 / -4.3 | -14.4 / -2.2 |
| Length | 5.18s | 5.18s | 5.04s | 5.04s | — | 8.00s | 4.92s |
| Frame | 768×1344 | 768×1344 | 1072×1928 | 720×1280 | — | 720×1280 | 576×1024 |
| Wait | 7s | 2:33 | 3:15 | 5:27 | — | 35s | 24s |
| Attempts | 1 | 1 | 1 | 1 | — | 2 | 1 |
| Cost | $0.40 | $0.30 | $0.98 | $0.50 | — | $3.20 | $0.22 |
The read
The easiest input on the sheet, and most models treated it that way. H3 Max, H3 reference-to-video, Kling and Veo all held the face, said the line, and picked up the mug on cue; Veo's is the most natural performance of the four and cost four times as much for three seconds it was not asked for. Wan darkened the room and re-lit her from the first frame, then drank. LTX re-framed to a tight close-up, dropped the mug and the table, and still passed the transcript check — the line was heard, the shot was not the shot. Seedance refused the frame as a possible likeness of a real person, and will refuse any photoreal face you give it.
One person's judgment from the takes above, 2026-09-03. The table is the evidence; this is the opinion.
what whisper heard
- H3 Max: I told you it would work. I just didn't say when.
- H3 ref2v: I told you it would work. I just didn't say when.
- Kling v3 Pro: I told you it would work, I just didn't say when.
- Wan 2.7: I told you it would work. I just didn't say when.
- Veo 3.1: I told you it would work. I just didn't say when.
- LTX-2: I told you it would work. I just didn't say when.