Model Bakeoffrendered on fal · 2026-09-03

One face, one line

A single human face has to hold its identity while the mouth matches a line the model voices itself.

establishing frame
the establishing frame, generated with Flux
The woman leans forward slightly, looks into the camera and says: "I told you it would work. I just didn't say when." Then she picks up the mug and takes a sip. The camera holds steady. Natural morning light, room tone, no music.
2 models received different prompt text — show each

Each model is given the prompt its input contract calls for: a multi-reference model is told which reference is which character, a segmented model gets shot segments. The frame, the references and the requested length are identical throughout. This is the honest version of "the same prompt".

H3 Max, H3 ref2v, Wan 2.7, Veo 3.1, LTX-2

The woman leans forward slightly, looks into the camera and says: "I told you it would work. I just didn't say when." Then she picks up the mug and takes a sip. The camera holds steady. Natural morning light, room tone, no music.

Kling v3 Pro

[{"prompt":"The woman leans forward slightly, looks into the camera and says: \"I told you it would work. I just didn't say when.\" Then she picks up the mug and takes a sip. The camera holds steady. Natural morning light, room tone, no music.","duration":"5"}]
  • “I told you it would work. I just didn't say when.”
5 s asked9:16Dialogue 86b7da01c0e8
0.0 ssound: off — click a take
H3 Max $0.40 · 7s · 768×1344 lines 1/1 · motion 1.67
H3 ref2v $0.30 · 2:33 · 768×1344 lines 1/1 · motion 1.79
Kling v3 Pro $0.98 · 3:15 · 1072×1928 lines 1/1 · motion 2.51
Wan 2.7 $0.50 · 5:27 · 720×1280 lines 1/1 · motion 1.80 · clips
Seedance 2.0 refused
The images or videos provided may contain likenesses of real people or other private information that cannot be processed."
Seedance 2.0no charge
Veo 3.1 $3.20 · 35s · 720×1280 lines 1/1 · motion 2.15 · 8.00s
LTX-2 $0.22 · 24s · 576×1024 lines 1/1 · motion 2.13
H3 MaxH3 ref2vKling v3 ProWan 2.7Seedance 2.0Veo 3.1LTX-2
Lines heard1/11/11/11/11/11/1
Voicemodel's ownmodel's ownmodel's ownmodel's ownmodel's ownmodel's own
Motion1.671.792.511.802.152.13
Scene changes000000
Loudness / peak-20.5 / -9.4-13.5 / -1.9-20.4 / -8-12.5 / -0.6-19.4 / -4.3-14.4 / -2.2
Length5.18s5.18s5.04s5.04s8.00s4.92s
Frame768×1344768×13441072×1928720×1280720×1280576×1024
Wait7s2:333:155:2735s24s
Attempts111121
Cost$0.40$0.30$0.98$0.50$3.20$0.22

The read

The easiest input on the sheet, and most models treated it that way. H3 Max, H3 reference-to-video, Kling and Veo all held the face, said the line, and picked up the mug on cue; Veo's is the most natural performance of the four and cost four times as much for three seconds it was not asked for. Wan darkened the room and re-lit her from the first frame, then drank. LTX re-framed to a tight close-up, dropped the mug and the table, and still passed the transcript check — the line was heard, the shot was not the shot. Seedance refused the frame as a possible likeness of a real person, and will refuse any photoreal face you give it.

One person's judgment from the takes above, 2026-09-03. The table is the evidence; this is the opinion.

what whisper heard