Models › H3 ref2v
MiniMax H3, reference-to-video
minimax/h3/reference-to-video
≤9 images≤3 video refs≤3 audio refs5–15 s768p native
price$0.06/s
wait p50 / max3:34 / 10:24
checks all green3 of 8
lines heard100%
refused0
spent here$3.32
The multi-reference sibling, and the only model here that embeds a supplied recording rather than imitating it. Re-stages the scene from the references, so adherence is better and continuity with the frame is worse.
Everything it rendered
8 inputs, 8 takes. Click any one for the full request and its measurements.
Identity across calls
| Character | Mean similarity | Appearances | Detail |
|---|---|---|---|
| The baker | 0.972 | 2 | 2 calls compared to their own centroid · 10/10 sampled frames passed the fidelity gate |
| serial-flyer | — | 0 | the entity was not found in any sampled frame — no score possible |
| The paper fox | — | 1 | only one input produced a usable appearance, so there is nothing to compare across calls |







