Developer
MiniMax

MiniMax H3 (Hailuo 3.0)

MiniMax's omni-modal video model, announced July 31, 2026: 33B open weights, 2K video up to 15 seconds, with stereo audio generated jointly.

Updated

Released
Jul 31, 2026
Type
Video Generation Model
License
MiniMax H3 Community License (open weights, jurisdiction-restricted)
On this page

Overview

MiniMax H3 is MiniMax's current flagship video generation model, announced on July 31, 2026 and shipped to consumers as Hailuo 3.0. MiniMax's release notes describe it as "a new-generation open general-purpose multimodal video model" that "understands creative intent across multimodal context — text, image, video, and audio." The API model ID is MiniMax-H3; third-party routers list it as minimax/hailuo-3; the marketing name in the Hailuo apps is Hailuo 3.0. All three refer to the same 33-billion-parameter model.

The output is what changed. H3 produces 4 to 15 seconds of video at up to 2K, 24 fps, with stereo audio generated in the same pass — dialogue, foley, room tone and music, not a soundtrack attached afterwards. Its predecessor, Hailuo 2.3, generates silent video capped at 1080p for six seconds. That is a generational gap, and it arrived after nine months in which MiniMax shipped language, speech and music models and no video model at all.

The architectural claim is the more interesting half. Earlier Hailuo releases were a menu of specialists — text-to-video, image-to-video, first-and-last-frame, subject reference, separate models for speech, effects and music — each with its own endpoint and its own constraints. H3 collapses those into one pretraining objective, with reference and editing relationships expressed in natural language rather than selected from a fixed task list. MiniMax names four components behind this: H3-VAE (a rebuilt tokenizer claiming a 4× gain in effective sequence length), the H3-Omni Transformer (separating understanding from generation workloads, for a claimed ~30% training-throughput gain), in-context regeneration (the base model regenerates its own low-resolution output at 2K instead of using a super-resolution module), and a contextual captioning pipeline distilling ~100K tokens down to ~4K. MiniMax has published no technical report and no benchmark table, so every one of those numbers is a vendor claim. Our launch write-up covers the announcement in detail: MiniMax H3: 2K Video With Native Audio, Weights Promised.

The weights landed, with pieces missing

At announcement the weights were a promise — "in the coming days, subject to applicable laws and regulations." They shipped on August 3, 2026 to MiniMaxAI/MiniMax-H3 on Hugging Face: two base checkpoints, FL2VA (text- and first/last-frame-to-audio-video) and Ref2VA (reference-conditioned), 33B parameters, under the MiniMax H3 Community License.

Two qualifiers matter more than the headline. First, H3-Context-IR and H3-Regenerate-2K were not released, so a local install generates at a 768-pixel short edge — the 2K path stays behind MiniMax's API. Second, the license excludes local-deployment rights in the United States, the European Union, the United Kingdom and the Republic of Korea without separate written authorization, and gates organizations above roughly $20M in annual revenue everywhere. "Open weights" is accurate; "open source, use it anywhere" is not. Read the license before you plan a deployment on it.

Capabilities

  • Omni-modal input in one prompt: Text, images, video clips and audio in a single request — up to nine images, three video clips and three audio clips, twelve files and 64 MB total, with a 7,000-character prompt. Relationships between them are described in language ("match the camera move in clip 1, have the character in image 3 sing this track").
  • Native stereo audio: Voice, sound effects, ambience and music generated jointly with the picture. No separate audio model, no post sync.
  • 2K output at 15 seconds: A 1440-pixel short edge at 24 fps, reached by in-context regeneration rather than an upscaler that never sees the reference material.
  • Instruction-based editing: Image-to-image, image-to-video, audio-to-audio and audio-video-to-audio-video reference and editing, expressed as natural-language instructions rather than a fixed task selection.
  • Multi-shot modelling: MiniMax lists coherent multi-shot generation as a capability of the unified architecture.
  • Text and brand rendering: MiniMax highlights legible on-screen text and accurate brand marks — the failure mode that keeps most video models out of advertising work.
  • Video-to-video motion transfer: Referencing camera movement or motion from a supplied clip.

Technical Specifications

  • API model IDs: MiniMax-H3, MiniMax-H3-Max
  • Parameters: 33B (dense omni-modal transformer)
  • Modes (H3): Text-to-video, image-to-video with first/last-frame roles, reference-to-video (images, video, audio). Image-to-video and reference-to-video are mutually exclusive in one request.
  • Modes (H3-Max): Text-to-video and image-to-video only; reference generation listed as coming soon
  • Resolutions: 768P or 2K (H3); 480P or 768P (H3-Max)
  • Duration: 4–15s (H3), 5–15s (H3-Max), integers only, 24 fps
  • Aspect ratios: 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, adaptive
  • Audio: Native stereo, generated with the video
  • Weights: Open — MiniMaxAI/MiniMax-H3, FL2VA and Ref2VA checkpoints; H3-Context-IR and H3-Regenerate-2K withheld; local ceiling 768-pixel short edge
  • License: MiniMax H3 Community License — jurisdiction-restricted, revenue-gated
  • API contract: Asynchronous task creation with polling or callback
  • Access: MiniMax Open Platform API, Hailuo AI web and mobile apps, third-party routers and hosts

MiniMax has published no model card with architecture depth, training-data summary or evaluation protocol. The parameter count is the only structural figure available.

Previous generation: Hailuo 2.3

Hailuo 2.3 shipped on October 28, 2025 and remains callable. It is closed-weight, silent, and available in two variants.

  • API model IDs: MiniMax-Hailuo-2.3, MiniMax-Hailuo-2.3-Fast
  • Modes: Text-to-video and image-to-video (2.3); image-to-video only (2.3-Fast)
  • Resolutions and durations: 768P at 6 or 10 seconds; 1080P at 6 seconds only. No 512P — that stays on MiniMax-Hailuo-02.
  • Audio: None. No audio parameter exists in either API reference.
  • Weights: Closed. No parameter count, architecture or training-data disclosure.
  • Still the only Hailuo models sold through video-point packages — H3 is pay-as-you-go or custom contract.

MiniMax's stated focus at the 2.3 launch was physical motion and micro-expression, "near-photorealistic visual effects in lighting direction, shadow transitions, and color tones," and stylized rendering for anime, illustration, ink wash and game CG. Those were vendor claims then and remain unbenchmarked now. Two older endpoints never moved: first-and-last-frame generation accepts only MiniMax-Hailuo-02, and subject-reference generation accepts only S2V-01.

Use Cases

  • Ad and product video with sound: The showcased case, and the one the pricing supports — a 6-second 2K clip with audio at $0.78 is cheap enough to iterate on, and legible on-screen text is a stated design target.
  • Reference-driven character work: Nine images, three clips and three audio tracks in one prompt is enough to hold a character, a location and a voice consistent across generations without a fine-tune.
  • Instruction-based editing passes: Re-cutting or restyling existing footage by describing the change, rather than re-running a task-specific endpoint.
  • Music and dialogue-led shorts: Joint audio-video means lip movement and vocals come out of the same pass; separate models require sync work in post.
  • On-premise 768p generation: Where the license permits it, the open checkpoints run locally — useful for content that cannot leave the building. The 2K stage does not come with them.
  • Cheap high-volume silent clips: Still Hailuo 2.3 Fast's job at $0.19 per 768P 6-second clip, if you need no audio and only image-to-video.

Performance / Benchmarks

MiniMax publishes no benchmark table for H3 — no arena Elo, no comparison chart, no technical report. Everything below is third-party, and the two arenas use different Elo scales that must not be compared to each other.

arena.ai Video Arena — September 4, 2026

BoardModelRankEloVotes
Image-to-videominimax-h31 of 471497±636,137
Image-to-videohailuo-2.3281261±5372,163
Text-to-videominimax-h38 of 481462±107,648
Text-to-videohailuo-2.3321205±682,971

H3 sits ahead of gemini-omni-1.1-flash (1488), wan3.0 (1481) and both Dreamina Seedance entries (1478 and 1477) on the image-to-video board, though the gap to second place is 9 Elo points against a ±6 confidence interval — a lead, not a separation. Its text-to-video placing is materially weaker: rank 8, roughly 50 points behind the Gemini Omni Flash entries. The vote counts are the caveat on both — 36,137 votes against Hailuo 2.3's 372,163, on a board carrying 1.9M votes total. These are early numbers.

Artificial Analysis Video Arena — August 2026

On the image-to-video board with audio, MiniMax H3 Max (post-trained by fal) ranks first at 1201 Elo (±11 over 2,177 samples, as of August 25, 2026), Dreamina Seedance 2.0 720p second at 1191, and MiniMax H3 third at 1187. Without audio, MiniMax H3 is reported at 1347 Elo, behind Wan 3.0 at 1358 — leading the open-weight field but not the board overall.

The honest reading: H3 is genuinely at the frontier on image-to-video and in the top ten on text-to-video, a jump of roughly 240 Elo over Hailuo 2.3 on the same board. It is the first open-weight model to top a major video arena. It is not a clear winner over the closed frontier, and the sample sizes behind its rankings are an order of magnitude smaller than the ones behind the incumbents.

Limitations

  • The open weights cannot legally be run locally in the US, EU, UK or South Korea without separate authorization from MiniMax, and organizations above roughly $20M revenue need authorization anywhere. For a large share of the audience for an open video model, this is a hosted-API model.
  • 2K is API-only. H3-Regenerate-2K was withheld from the release, so a local install caps out at a 768-pixel short edge. The headline capability is not in the download.
  • No technical report, no model card, no vendor benchmarks. The 4× tokenizer gain, the ~30% throughput improvement and every quality claim are unverified vendor statements.
  • 15 seconds is the ceiling. Seedance 2.5 advertises 30-second single-pass generation; H3 does not go past 15.
  • No video-point packages. Volume buyers on MiniMax's package tiers cannot spend points on H3 — it is pay-as-you-go or a sales conversation.
  • Reference and image-to-video are mutually exclusive per request, which forces a choice between first/last-frame control and multimodal referencing.
  • 768P on H3 is documented as closed beta by some third-party write-ups even though MiniMax's rate card lists a price for it — confirm access before designing around the cheaper tier.
  • The legacy line is frozen. Hailuo 2.3 is silent, capped at 1080p/6s, and two of MiniMax's video endpoints still run on models from 2025.

Pricing & Access

MiniMax H3 — pay as you go

Per second of output, in USD, from MiniMax's rate card.

ModelOutputPrice
MiniMax-H32K$0.13/s ($7.80/min)
MiniMax-H3768P$0.08/s ($4.80/min)
MiniMax-H3-Max768P$0.08/s
MiniMax-H3-Max480P$0.05/s

Input materials for H3: audio free, first five reference images free then $0.04 each, reference video billed by input duration at the output rate. H3-Max input materials are not billed "for now." H3-Context-IR is billed separately as a token task at $0.90/M input and $3.60/M output. A 6-second 2K clip with audio costs $0.78.

For context, Artificial Analysis's cost comparison puts Kling 3.0 at roughly $20.16 per minute at 1080p and Dreamina Seedance 2.0 at roughly $22.45 — H3 at 2K is around a third of either. Those competitor figures are Artificial Analysis's derivation, not vendor rate cards.

Hailuo 2.3 — pay as you go

Per video, not per second.

ModelOutputPrice per video
MiniMax-Hailuo-2.3-Fast768P, 6s$0.19
MiniMax-Hailuo-2.3-Fast768P, 10s$0.32
MiniMax-Hailuo-2.3-Fast1080P, 6s$0.33
MiniMax-Hailuo-2.3768P, 6s$0.28
MiniMax-Hailuo-2.3768P, 10s$0.56
MiniMax-Hailuo-2.31080P, 6s$0.49

Failed generations are not billed: "video generation failures or videos that trigger security review will not result in a deduction."

Video packages

Standard $1,000/mo (3,760 video points, 20 RPM), Pro $2,500 (9,920, 30), Scale $4,500 (18,900, 40), Business $6,000 (26,780, 50), plus a custom tier. Deductions: MiniMax-Hailuo-2.3-Fast 0.7–1.3 points, MiniMax-Hailuo-2.3 and MiniMax-Hailuo-02 1–2 points. H3 is excluded — MiniMax directs H3 users to pay-as-you-go or a custom plan.

Consumer

Hailuo AI subscriptions, from its payment policy: Standard $14.99/mo (1,000 credits), Pro $54.99/mo (4,500), Master $119.99/mo (10,000), Ultra $124.99/mo (12,000), Max $199.99/mo (20,000). Credits expire monthly and do not roll over. A legacy Unlimited plan at $94.99/mo, purchasable only before June 18, 2025, is restricted to "the Hailuo01 model."

Ecosystem & Tools

  • MiniMax Open Platform — API keys, release notes, rate cards
  • Video generation guide — the H3 and H3-Max mode and resolution matrix
  • Image-to-video API reference — the legacy model enum, which still accepts both 2.3 variants
  • MiniMaxAI/MiniMax-H3 — the open FL2VA and Ref2VA checkpoints and the Community License
  • Hailuo AI — the consumer web and mobile apps, branded Hailuo 3.0
  • fal — hosted H3 Max, which fal says its research team post-trained
  • Replicate — hosted Hailuo 2.3, 768p and 1080p
  • MiniMax — the vendor's product surface in our tools catalog
  • ComfyUI — community nodes for the open checkpoints landed alongside the August 3 weight release

Community & Resources

Frequently Asked Questions

Yes. MiniMax announced H3 on July 31, 2026 and it ships as Hailuo 3.0 in the consumer apps. MiniMax's API release notes carry the entry "Jul. 31, 2026 — MiniMax H3: a new-generation open general-purpose multimodal video model," the API model ID is MiniMax-H3, and OpenRouter routes it as minimax/hailuo-3. An earlier version of this page said no such model existed; that was true until July 31, 2026 and is not true now.
A 33-billion-parameter omni-modal generation model. It reads text, images, video and audio as a single context and returns 4–15 second video at up to 2K, 24 fps, with stereo audio generated jointly with the picture rather than dubbed on afterwards. It also does instruction-based editing, described in natural language instead of selected from a fixed task menu.
The weights shipped on August 3, 2026 to MiniMaxAI/MiniMax-H3 on Hugging Face — two base checkpoints, FL2VA and Ref2VA — under the MiniMax H3 Community License. Two pipeline pieces were withheld: H3-Context-IR (prompt preprocessing) and H3-Regenerate-2K (the 2K stage), so a local install tops out at a 768-pixel short edge.
Not under the license as written. The MiniMax H3 Community License excludes the United States, the European Union, the United Kingdom and the Republic of Korea from local-deployment rights without separate written authorization, and organizations above roughly $20M in annual revenue need written authorization regardless of location. The hosted API is not geographically restricted this way. Read the license yourself before deploying; this is a summary, not legal advice.
MiniMax's pay-as-you-go rate card prices H3 output per second: $0.13/second at 2K ($7.80 per minute) and $0.08/second at 768P. Reference images are free for the first five and $0.04 each after; reference audio is free; reference video is billed at the output rate. A 6-second 2K clip with audio is $0.78. H3 is not sold through MiniMax's video-point packages — pay-as-you-go or a custom contract only.
On the published comparison, yes. H3 at 2K works out to roughly $7.80 per minute against roughly $20.16 per minute for Kling 3.0 at 1080p and roughly $22.45 for Dreamina Seedance 2.0 at 1080p. Those two figures are derived from Artificial Analysis's cost comparison rather than from Kuaishou's or ByteDance's own rate cards, and third-party hosts charge their own margins on top.
Yes. Stereo audio — dialogue, foley, ambience and music — is generated jointly with the video in the same pass. This is the clearest break from Hailuo 2.3, which is silent and exposes no audio parameter anywhere in its API.
768P or 2K (a 1440-pixel short edge) at 24 fps, 4 to 15 seconds, in 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16. A single request accepts up to nine reference images, three reference video clips and three reference audio clips, twelve files in total.
A speed-optimized H3 variant. MiniMax's rate card lists MiniMax-H3-Max at $0.05/second for 480P and $0.08/second for 768P, text-to-video and image-to-video only, with reference generation listed as coming soon and input materials not billed for now. fal separately markets a "MiniMax H3 Max" that it says its own research team post-trained; MiniMax does not describe the provenance of its H3-Max rows, so treat the two as related but not confirmed identical.
Yes, as the previous generation. MiniMax's image-to-video API reference still accepts MiniMax-Hailuo-2.3 and MiniMax-Hailuo-2.3-Fast, the per-video rate card still lists them ($0.28 for 768P/6s, $0.49 for 1080P/6s), and they remain the only Hailuo models sold through the video-point packages. They generate silent video and are capped at 1080p for 6 seconds.
On arena.ai's image-to-video board (September 4, 2026) minimax-h3 is rank 1 of 47 at 1497±6 Elo over 36,137 votes; on text-to-video it is rank 8 of 48 at 1462±10. On Artificial Analysis's image-to-video arena with audio, MiniMax H3 Max leads at 1201 Elo and MiniMax H3 is third at 1187. The two arenas use different Elo scales and are not comparable to each other.
MiniMax, a Shanghai-based AI company that listed on the Hong Kong Stock Exchange on January 9, 2026. Global Times reported the stock opened at HK$235.40 against an IPO price of HK$165 before "closing 109.1 percent higher," pushing the valuation past HK$103 billion (US$13.2 billion). MiniMax also builds the MiniMax-M3 language model.

Explore More Models

Discover other AI models and compare their capabilities.