The same fox, a different scene
The identity test: the same paper fox, a different setting and action, rendered in a separate call. Nothing carries over between the two except the reference images.
Same character, separate call: Stylised animation · how consistent each model was
| H3 Max | LTX-2 | |
|---|---|---|
| Lines heard | — | — |
| Voice | model's own | model's own |
| Motion | 0.75 | 3.52 |
| Scene changes | 0 | 0 |
| Loudness / peak | -32.2 / -18.1 | -54.9 / -29.9 |
| Length | 5.18s | 4.92s |
| Frame | 768×1344 | 576×1024 |
| Wait | 8s | 45s |
| Attempts | 1 | 1 |
| Cost | $0.40 | $0.22 |
The read
The identity test, and the honest answer is that five seconds is not long enough to break most models. H3 Max and LTX both produced a recognisable fox at the stream from the same reference set that produced the fox in the forest, and the embedding check scores them 0.96 and 0.94 against their own centroids. What the numbers do not capture is that H3 Max kept the paper texture and the layered cut edges while LTX flattened them slightly and moved the camera more than the fox. That is the same split the forest shot showed. The useful finding here is not the ranking, it is that a second scene from the same references is cheap to test and almost nobody does it.
One person's judgment from the takes above, 2026-09-03. The table is the evidence; this is the opinion.
what whisper heard
- H3 Max: (water splashing)