MiniMax AI video models
MiniMax H3 models support different ways to direct a shot: text, designed frames, and multimodal references. The key distinction is whether you need mixed reference media or a simpler text-and-frame workflow. Compare those requirements before evaluating resolution and credit cost.
3 models available in Starpop. Updated .
Compare MiniMax video models
| Model / best use | Inputs | Output options | Estimated cost |
|---|---|---|---|
| MiniMaxvideoMiniMax H3 Native-audio cinematic video | Text, Image, Video, Audio | 480p, 768p, 2K, 4K5–15s | $0.05–$0.16 USD / second0.5–1.6 credits / second |
| MiniMaxvideoMiniMax H3 Max Prompt-faithful cinematic video | Text, Image, Video, Audio | 480p, 768p, 1080p5–15s | $0.05–$0.16 USD / second0.5–1.6 credits / second |
| MiniMaxvideoMiniMax H3 Max Turbo Lower-cost video iteration from text | Text, Image | 480p, 768p, 1080p5–15s | $0.025–$0.08 USD / second0.25–0.8 credits / second |
Choosing a MiniMax model
Working with mixed references
Compare H3 and H3 Max when the brief combines image, video, and audio references. Give each reference a purpose, and check file-count and duration limits in the model guide before preparing the full input set.
Iterating from text or frames
H3 Max Turbo suits text-and-frame variations at a lower credit cost than H3 Max. Choose H3 Max when the shot requires mixed image, video, and audio reference inputs.
Checking the delivery format
Frame-based generation can follow the source image's canvas. Review aspect ratio and the model's resolution notes before generating, especially when a higher-resolution option uses refinement of a lower-resolution source.