Model Bakeoffrendered on fal · 2026-09-03

The same fox, a different scene

The identity test: the same paper fox, a different setting and action, rendered in a separate call. Nothing carries over between the two except the reference images.

establishing frame
the establishing frame, generated with Flux
The paper fox leans down to the paper stream and drinks, its ears flicking back, then lifts its head and looks off to the right. The layered paper water ripples beneath it. Keep the paper cut-out look throughout: visible texture, cut edges, layered shadows. No smooth 3D shading, no photorealism. Gentle water and paper rustling, no music.
5 s asked9:16Stylised character: The paper foxe40c0e3ff9fc

Same character, separate call: Stylised animation · how consistent each model was

0.0 ssound: off — click a take
H3 Max $0.40 · 8s · 768×1344 motion 0.75
LTX-2 $0.22 · 45s · 576×1024 motion 3.52
H3 MaxLTX-2
Lines heard
Voicemodel's ownmodel's own
Motion0.753.52
Scene changes00
Loudness / peak-32.2 / -18.1-54.9 / -29.9
Length5.18s4.92s
Frame768×1344576×1024
Wait8s45s
Attempts11
Cost$0.40$0.22

The read

The identity test, and the honest answer is that five seconds is not long enough to break most models. H3 Max and LTX both produced a recognisable fox at the stream from the same reference set that produced the fox in the forest, and the embedding check scores them 0.96 and 0.94 against their own centroids. What the numbers do not capture is that H3 Max kept the paper texture and the layered cut edges while LTX flattened them slightly and moved the camera more than the fox. That is the same split the forest shot showed. The useful finding here is not the ranking, it is that a second scene from the same references is cheap to test and almost nobody does it.

One person's judgment from the takes above, 2026-09-03. The table is the evidence; this is the opinion.

what whisper heard