Model Highlights: MiniMax H3

MiniMax H3 is a powerful new 33B parameter omni-modal generative system designed for unified understanding and generation of text, images, video, and audio. It is capable of generating highly synchronized videos with native 32kHz stereo audio, supporting up to 15 seconds of runtime and up to 2K resolution.

More info is available on AImodels.fyi here!

Key Highlights

  • Omni-Modal Inputs: Seamlessly processes free-form multimodal prompts containing text, images (up to 9), reference videos, and reference audio to drive highly specific generations.

  • Three-Tiered Architecture:

    • H3-Base (Open Source): The core 33B-parameter dense, single-stream Omni-Transformer that generates 768p video and audio.

    • H3-Context-IR (API): A hosted orchestration module that deeply understands and enhances complex multimodal prompts before generation.

    • H3-Regenerate-2K (API): A novel in-context module that feeds the base 768p output back into the model to natively regenerate it at high-fidelity 2K resolution.

  • Two Specialized Checkpoints: Available in FL2VA (First/Last-frame to Audio-Video) for standard text/image-to-video, and Ref2VA (Reference to Audio-Video) for advanced editing, voice timbre referencing, and style mimicry.

  • Advanced Encoding: Leverages the full pretrained weights of Qwen3-VL-32B for text/visual encoding, paired with highly optimized, temporally causal Visual and Audio VAEs.

  • Broad Deployment Support: The H3-Base checkpoints are ready for local and cloud deployment out-of-the-box with SGLang, vLLM, Diffusers, and ComfyUI.

License & Availability: The core H3-Base model weights are available to download and deploy locally under the MiniMax H3 Community License Agreement. The supplementary H3-Context-IR and H3-Regenerate-2K modules are currently available via the MiniMax Open Platform API.

Scroll to Top