MiniMax H3 is a general-purpose multimodal generation model created by MiniMax. It accepts creative context across text, images, video, and audio, understands the relationships among those inputs, and generates video with synchronized native stereo sound. MiniMax announced output up to 15 seconds at 2K resolution, positioning H3 for advertising, branding, ecommerce, film, product design, UI, gaming, and other commercial creative work.
Unlike video systems divided into separate text-to-video, image-to-video, motion-reference, subject-reference, audio, and editing models, H3 is designed to express those tasks through natural-language instructions inside one generalized system. Its supported workflows include text-to-audio-video, first-and-last-frame generation, reference-to-audio-video, multimodal editing, motion transfer, native multi-shot modeling, and joint generation of voice, sound effects, and music.
MiniMax released H3-Base weights under the MiniMax H3 Community License. The official model card describes a 33B dense H3-Omni Transformer, H3 visual and audio VAEs, and separate checkpoints for first-or-last-frame and reference-to-audio-video tasks. This VisionStory page is an independent model profile: MiniMax built H3, while VisionStory helps creators explore AI video workflows and compare leading video models.








