01

What MiniMax H3 is

MiniMax describes H3 as a general-purpose omni-modal video generation system. The August 2026 announcement says it can interpret combinations of text, images, video, and audio and generate video with native stereo audio. Published output options include 4–15 second clips, 24 frames per second, multiple aspect ratios, a default shorter side of 768 pixels, and an optional 2K regeneration workflow.

02

Two base variants serve different inputs

H3-Base-FL2VA covers text-to-video plus first-frame, last-frame, or first-and-last-frame conditioning. H3-Base-Ref2VA accepts broader reference packages: the release describes up to nine images, up to three short video clips, and up to three audio clips, with a maximum of twelve files across supported input types. Audio cannot be the only reference input. Those limits describe the reviewed release and should be checked again before building a production uploader.

  • FL2VA: begin from text and optionally one or two boundary frames.
  • Ref2VA: combine visual, motion, and audio references for a more constrained result.
  • Both base checkpoints generate audio and video; input choice determines the workflow, not a quality ranking.
03

The open release is not the entire hosted pipeline

The published system has three major parts. H3-Context-IR interprets and restructures complex multimodal instructions. H3-Base generates a 768p audio-video result. H3-Regenerate-2K uses the base result plus original context to regenerate at higher resolution. MiniMax states that the base checkpoints are open, while Context-IR remains a hosted API and the 2K regeneration module was not open-sourced at the review date.

04

Local deployment starts with hardware and workflow constraints

The H3-Omni-Transformer is described as a 33-billion-parameter dense model, and the released checkpoints use BF16. That makes local inference a substantial infrastructure project rather than a typical laptop installation. MiniMax documents SGLang, vLLM, diffusers, and ComfyUI paths, but teams still need to validate memory requirements, storage, runtime compatibility, download size, generation time, and output quality on their own hardware.

A defensible pilot measures one exact target clip on the intended stack. Record peak memory, total generation time, failed runs, moderation behavior, and the human time required to prepare references. Do not infer these values from parameter count alone.

05

Choose local base, hybrid 2K, or hosted generation deliberately

Use the local base path when control over the checkpoint and a 768p validation workflow matters more than convenience. Use the documented hybrid workflow when you can accept hosted Context-IR and regeneration calls around a local base. Use the fully hosted API when operational simplicity matters more than running the base weights yourself. “Local” and “private” are not synonyms if prompts, references, or outputs are still sent to hosted components.

06

Verify before continuing

Which files leave your environment? Is 768p base output sufficient, or does the workflow depend on hosted 2K regeneration? Have you reviewed the current community license and acceptable-use terms? Can your hardware run the selected checkpoint at the intended duration? Will reference media rights and consent survive the generation workflow?

  • Which files leave your environment?
  • Is 768p base output sufficient, or does the workflow depend on hosted 2K regeneration?
  • Have you reviewed the current community license and acceptable-use terms?
  • Can your hardware run the selected checkpoint at the intended duration?
  • Will reference media rights and consent survive the generation workflow?

QUESTIONS THIS ANSWERS

Questions this answers

  • What is MiniMax H3 and which parts are open source?
  • Can MiniMax H3 generate video with audio?

Found something wrong? Report an error or read the corrections policy.

2 SOURCESEvidence ledger

Sources

  1. 01
  2. 02
    Video generation guide ↗

    MiniMax · accessed 9 Sept 2026