Developer
MiniMax

MiniMax-M3

MiniMax's multimodal flagship, released June 1, 2026. A 428B-parameter MoE with ~23B active, a 1M-token context, and open weights under a community license.

Updated

Released
Jun 1, 2026
Type
Multimodal Language Model
Context window
1M tokens
Pricing
$0.30 / $1.20 per Mtok
License
MiniMax Community License (open weights, restricted)
On this page

Overview

MiniMax-M3 is MiniMax's native-multimodal flagship, released on June 1, 2026. MiniMax's own API release notes date it "Jun. 1, 2026," describing a model for "agentic reasoning, tool use, coding, multimodal chat input, and long-context tasks." The Hugging Face model card summarizes it in one line: "MiniMax-M3 is a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters."

The architectural bet is sparse attention. M3 is the first MiniMax flagship built on MiniMax Sparse Attention (MSA), which the model card calls "a high-performance sparse attention operator designed for million-token contexts." The launch blog states the payoff plainly: "at a context length of 1 million, M3's per-token compute is just 1/20 that of the previous-generation model," with "more than 9ร— in the prefilling stage and more than 15x in the decoding stage." The technical report, arXiv:2606.13392, landed June 11, 2026 โ€” ten days after the model.

The second bet is native multimodality. Rather than bolting a vision encoder onto a finished text model, MiniMax states that "M3 is a model that has undergone mixed-modality training from Step 0." Hugging Face classifies the repository as image-text-to-text; it accepts text, image, and video, and emits text.

M3 sits at the end of a fast release cadence documented in MiniMax's release notes: MiniMax-Text-01 (Jan 15, 2025) โ†’ M2 (Oct 27, 2025) โ†’ M2.1 (Dec 22, 2025) โ†’ M2.5 (Feb 2026) โ†’ M2.7 (Mar 18, 2026) โ†’ M3 (Jun 1, 2026).

As of September 6, 2026 that line still ends at M3: MiniMax's release notes record no M-series language model after it, only generative-media releases โ€” a video model on July 31, 2026 and MiniMax Music 3.0 on August 13. The video line is documented separately.

A note on the license

MiniMax-M3 is not MIT licensed, and the "open source" framing does not survive contact with the license text.

MiniMax's launch blog promised to "open-source the corresponding model weights." What shipped is the MINIMAX COMMUNITY LICENSE. Hugging Face tags the repository license: other with license_name: minimax-community. The base grant covers use "for non-commercial purposes"; any Commercial Use triggers two conditions. The license requires that "you shall prominently display 'Built with MiniMax M3' on a related website, user interface, blogpost, about page or product documentation," and that organizations whose products "generate more than 20 million US dollars (or equivalent in other currencies) in yearly revenue" must "obtain a separate, prior written authorization from MiniMax by contacting api@minimax.io" โ€” below that threshold a one-time notice to the same address suffices. It further prohibits military use, exploitation of minors, harmful misinformation, and hate speech. There are no geographic restrictions.

Under the definition used on this site, that is open weights under a restricted license โ€” not open source.

The regression did not begin with M3. MiniMax-M2 shipped under MIT in October 2025; today its Hugging Face card carries license: other with license_name: modified-mit, as does M2.5. Decrypt reported on April 13, 2026 that with M2.7 "commercial use now requires written authorization from MiniMax," quoting MiniMax developer-relations head Ryan Lee on degraded third-party hosting: "A fully permissive license meant we had no way to push back on any of that." M3 continues that policy under a new name.

Capabilities

  • Million-token context via MSA: max_position_embeddings in config.json is exactly 1,048,576. MSA is a blockwise sparse attention mechanism over a Grouped Query Attention backbone, not a compressed-latent scheme.
  • Native multimodal input: Text, image, and video in a single multimodal model trained mixed-modality from step zero, with a CLIP-style vision tower.
  • Sparse MoE efficiency: A 428B-parameter mixture of experts activating ~23B per token โ€” 128 routed experts, top-4 routing, plus one shared expert.
  • Three-way reasoning control: A thinking parameter accepting enabled, adaptive, or disabled, letting one deployment span deep chain-of-thought and low-latency turns.
  • Agentic tool use: MiniMax reports 74.2% on MCP Atlas, a benchmark of Model Context Protocol tool orchestration, and 66.0% on Terminal-Bench 2.1.
  • Long-horizon coding: MiniMax positions M3 as reaching "frontier-level performance across long-horizon agentic benchmarks, excelling in both coding and cowork."
  • Broad serving support: SGLang, vLLM, Transformers (minimax_m3_vl), KTransformers, and unsloth all ship documented recipes.

Technical Specifications

  • Hugging Face repository: MiniMaxAI/MiniMax-M3
  • Total parameters: ~428B per the model card; 427,040,140,160 exactly per Hugging Face's safetensors metadata
  • Active parameters: ~23B per token (model card); NVIDIA's deployment blog states 22B
  • Experts: 128 routed, top-4 activated per token, plus 1 shared expert (config.json)
  • Layers: 60 ยท Hidden size: 6,144 ยท Attention heads: 64 ยท KV heads: 4 (GQA)
  • Vocabulary: 200,064 tokens ยท Weights dtype: bfloat16
  • Vision tower: CLIP-style encoder, 32 layers, hidden size 1,280
  • Architecture class: MiniMaxM3SparseForConditionalGeneration (model_type: minimax_m3_vl)
  • Attention: MiniMax Sparse Attention (MSA), arXiv:2606.13392
  • Context window: 1,048,576 tokens
  • License: MiniMax Community License (license: other)
  • Recommended sampling: temperature=1.0, top_p=0.95
  • Reasoning control: thinking = enabled | adaptive | disabled

MiniMax does not publish a knowledge cutoff or a training-token count for M3, so this page does not state one.

Use Cases

  • Whole-repository code work: A 1M-token context window admits large codebases in one pass โ€” the workload MSA was built to make affordable.
  • Long-horizon agents: Agentic workflows that accumulate long tool-call transcripts benefit most from the claimed 1/20 per-token compute at 1M context.
  • Document and video understanding: Native video and image input, evaluated by MiniMax on Video-MME at 512 frames, suits multimodal archive and media analysis.
  • Tool-calling backends: The thinking: adaptive mode plus MCP-oriented evaluation targets function calling services that mix trivial and hard requests.
  • Cost-sensitive inference at scale: $0.30 / $1.20 per million tokens sits materially below most frontier-adjacent hosted APIs.
  • Self-hosting under compliance review: Weights are downloadable, but the $20M revenue authorization gate makes this a legal question first.

Performance / Benchmarks

Vendor-reported. Every figure below appears as text in MiniMax's own launch blog. The model card's benchmark comparison is published only as an image (figures/benchmark.jpeg) and the repository's Hugging Face model-index is null. The repo's .eval_results/minimax-m3.yaml, added in June 2026, states on its face that it was "Extracted from the model card benchmark graph" โ€” a transcription of that chart, not an audit of it.

BenchmarkMiniMax-M3 (vendor-reported)
SWE-Bench Pro59.0%
MCP Atlas74.2%
OSWorld-Verified70.06%
Terminal-Bench 2.166.0%
SWE-fficiency34.8%
KernelBench Hard28.8%
Video-MME (512 frames)84.6

SWE-Bench Pro, MCP Atlas, Terminal-Bench 2.1, SWE-fficiency, and KernelBench Hard appear in the blog's headline list; OSWorld-Verified 70.06% and the 84.6 Video-MME result at 512 frames appear only in its evaluation-methodology notes.

MiniMax's blog names SWE-Bench Verified, Claw-Eval, MMMU Pro, VideoMMMU, BrowseComp, PaperBench, Apex-Agents, IMO 2025, and USAMO 2026 without publishing M3 scores in text for them. The repo's .eval_results/minimax-m3.yaml transcribes four off the benchmark JPEG โ€” SWE-Bench Verified 80.5, MMMU-Pro 78.1, Claw-Eval 74.5, Apex-Agents 27.7 โ€” plus Video-MME (w/ sub) 85.4. As chart readings of MiniMax's own image, they carry no more weight than the vendor numbers above.

Third-party placements

Artificial Analysis reports 36 on the Artificial Analysis Intelligence Index v4.2, the version live on September 6, 2026 and an aggregate of ten evaluations, with output throughput of 84.6 tokens per second. On that same v4.2 scale the leaderboard shows Kimi K3 (max) at 50, GLM-5.3 (max) at 49, GLM-5.2 (max) at 43, DeepSeek V4 Pro 0813 (max) at 42, and MiMo-V2.5-Pro at 33.

The index version moved between v4.1 and v4.2, which matters more than the direction of the number: M3 read 44 on v4.1 and reads 36 on v4.2, with nothing about the deployed model changed in between. Compare within one index version, never across them. The widely circulated claim that M3 "scores 55" traces to Artificial Analysis's article of June 8, 2026 โ€” published before the weights shipped, naming no index version, and predating both v4.1 and v4.2. It is not a current score.

arena.ai (the former LMArena, rebranded January 28, 2026, and a different Elo scale from Artificial Analysis) now places M3 at rank 80 on Text Arena with 1443ยฑ4 Elo and rank 45 with 1487 (+7/โˆ’7) on Code Arena's WebDev board โ€” down from rank 54 (1447) and rank 16 (1501) in July. The text Elo barely moved, so what changed is the field around it. Note that arena.ai/leaderboard/code resolves to WebDev, a web-development board rather than a general coding one.

On the MSA speedup numbers

Two sets of speedup figures circulate and they measure different things. The model card claims "9ร— prefill and 15ร— decode speedups compared to M2 at 1M context, reducing per-token compute to 1/20." The arXiv paper reports, "on a 109B-parameter model with native multimodal training," a "28.4x" reduction in per-token attention compute at 1M context and "14.2x prefill and 7.6x decoding wall-clock speedups" on H800 GPUs. The paper's numbers are not M3's.

Limitations

  • Not open source. The MiniMax Community License gates commercial use above $20M annual revenue behind prior written authorization and mandates "Built with MiniMax M3" attribution. Redistribution and commercial deployment are not free of conditions.
  • A deliberate openness regression. M2 shipped under MIT in October 2025; M2.7 and M3 do not. Teams that adopted the M-series on permissive terms cannot assume continuity in future releases.
  • All headline benchmarks are vendor-reported. SWE-Bench Pro 59.0, MCP Atlas 74.2, and the rest come from MiniMax. This page found no independent replication; the model card ships its comparison chart as a JPEG, and the repo's .eval_results YAML only transcribes that chart.
  • Third-party scoring is far less flattering than the launch narrative. Artificial Analysis's live Intelligence Index v4.2 places M3 at 36, behind Kimi K3 (max) at 50, GLM-5.3 (max) at 49 and DeepSeek V4 Pro 0813 (max) at 42. Its arena.ai placements have slid to rank 80 (Text) and rank 45 (WebDev) as the field has filled in.
  • Long context is not free. Above 512K input tokens the API price doubles to $0.60 / $2.40 per million tokens, so the 1M window costs more than the headline rate implies.
  • 428B parameters is a serious self-hosting bill. At bfloat16 the weights alone are roughly 854 GB; sparse activation reduces compute per token, not the memory needed to hold the experts.
  • No published knowledge cutoff or training-data disclosure. Ground time-sensitive queries with retrieval or tool use.
  • Absolute agentic scores remain low. SWE-fficiency at 34.8% and KernelBench Hard at 28.8% โ€” MiniMax's own numbers โ€” show long-horizon autonomy is unsolved.

Pricing & Access

MiniMax publishes a USD rate card on its pay-as-you-go pricing page, verified September 6, 2026 and unchanged since launch. It labels the current rates "Permanent 50% off" list.

TierInput (per 1M)Output (per 1M)Cache read (per 1M)
Standard, โ‰ค 512K input tokens$0.30$1.20$0.06
Standard, > 512K input tokens$0.60$2.40$0.12
Priority, โ‰ค 512K input tokens$0.45$1.80$0.09
Priority, > 512K input tokens$0.90$3.60$0.18

Subscription token plans run Plus at $20/month (~1.7B tokens), Max at $50/month (~5.1B tokens), and Ultra at $120/month (~9.8B tokens).

Self-hosting: weights are downloadable from Hugging Face via hf download MiniMaxAI/MiniMax-M3, but see the license section โ€” this is not an unconditional grant.

Third-party hosting: OpenRouter lists minimax/minimax-m3 at 1.05M-token context and undercuts MiniMax's own rate card at $0.23 / $0.96 per million tokens, alongside :free and :batch (524K context) variants. As of September 6, 2026 OpenRouter's MiniMax catalogue carries 16 entries in total, of which the language models are M3, M2.7, M2.5, M2.1, M2, M2-her, M1 and MiniMax-01; the rest are video and speech models.

Adoption signal: Hugging Face reports 175,846 downloads over the trailing 30 days for MiniMaxAI/MiniMax-M3 โ€” down from 233,589 in July โ€” plus 288,782 for the quantised MiniMax-M3-MXFP8. The older, text-only MiniMax-M2.7 draws 1,222,347 over the same window. Three months after launch, M3's two official repos combined have not come close to displacing it.

Legacy models and what a migration off M2.5 actually involves

MiniMax's models guide splits the M-series in two: M3, M2.7 and M2.7-highspeed are listed as current; M2.5, M2.5-highspeed, M2.1, M2.1-highspeed and M2 are listed as legacy โ€” which on that page means "do not start a new integration here," not "switched off."

MiniMax has published no retirement date for M2.5. Checked on September 6, 2026, its release notes, models guide and pay-as-you-go pricing page carry no deprecation notice or end-of-service date for any M-series model, and M2.5 is still callable by exact model ID at $0.30 input / $1.20 output per million tokens, cache read $0.03.

The withdrawals that have happened are host-level and run on each host's own calendar. Fireworks removed MiniMax M2.5 from serverless on June 17, 2026, with the instruction "Migrate to MiniMax M2.7"; OpenRouter still routes minimax/minimax-m2.5 across eight providers at $0.27 / $0.95. That asymmetry is the planning problem: the binding date is your host's, not MiniMax's, and a model MiniMax still serves can vanish from the endpoint you actually call. The two landing spots are M2.7 for text and tool workflows at the same headline price, and M3 where 1M context or image and video input is the reason to move at all.

Ecosystem & Tools

Community & Resources

Frequently Asked Questions

June 1, 2026. MiniMax's own API release notes date the MiniMax M3 entry Jun. 1, 2026, and the GitHub repository was created on June 1, 2026. The weights came later: the Hugging Face repository's first commit is dated June 12, 2026, matching the launch blog's promise that "over the next 10 days, we will release the model's technical report and open-source the corresponding model weights."
No. This is the single most repeated error about M3. Hugging Face tags MiniMaxAI/MiniMax-M3 with license: other and license_name: minimax-community โ€” the MiniMax Community License. For any commercial use it requires you to "prominently display 'Built with MiniMax M3'", and organizations earning over $20 million USD annually from products built on it must "obtain a separate, prior written authorization from MiniMax" (below that threshold, a one-time notice suffices). That is open weights, not open source.
MiniMax's pay-as-you-go page lists $0.30 per million input tokens and $1.20 per million output tokens for requests with 512K or fewer input tokens, with prompt-cache reads at $0.06. Above 512K input tokens the rate doubles to $0.60 / $2.40. A priority tier costs 1.5x standard. MiniMax labels the current rates "Permanent 50% off" list.
MSA is the attention architecture behind M3's 1M-token context, documented in arXiv:2606.13392 ("MiniMax Sparse Attention", submitted June 11, 2026). The paper describes "a blockwise sparse attention built upon Grouped Query Attention" in which an Index Branch selects the relevant key-value blocks for each attention group, so attention runs over selected blocks rather than the full context.
The model card states "~428B parameters and ~23B activated parameters." Hugging Face's safetensors metadata puts the exact total at 427,040,140,160 parameters. NVIDIA's deployment blog describes it as a "428B parameter" model with "22B" active and "Total 128, 4 experts activated per token." MiniMax's config.json confirms 128 routed experts, top-4 routing, plus one shared expert.
No, and the live figure has moved twice. Artificial Analysis's model page reports 36 on Intelligence Index v4.2, the version live on September 6, 2026; earlier in the summer the same page read 44 on v4.1. The widely quoted 55 comes from Artificial Analysis's article of June 8, 2026, published a week after launch and before the weights shipped, and names no index version at all. Scores from different index versions are not comparable, so 36, 44 and 55 are three scales rather than a decline.
Yes, and natively so. Hugging Face's pipeline tag is image-text-to-text and the architecture is MiniMaxM3SparseForConditionalGeneration, with a CLIP-style vision tower in config.json. MiniMax says "M3 undergoes mixed-modality training from the very first step, enabling deeper semantic fusion across text, image, and video." It accepts text, image, and video input and returns text.
Through a thinking parameter with three values, per the model card: "enabled โ€” Reasoning is always enabled," "adaptive โ€” M3 automatically determines when additional reasoning is beneficial," and "disabled โ€” Reasoning is disabled to minimize latency and maximize throughput."
No. The 59.0% figure appears in MiniMax's own launch blog and is vendor-reported. This page has found no independent replication of it. The model card's full benchmark comparison is published only as an image (figures/benchmark.jpeg) and Hugging Face's model-index for the repository is null. A .eval_results/minimax-m3.yaml file was added to the repository in June 2026, but it is an explicit transcription of that image โ€” "Extracted from the model card benchmark graph" โ€” not an independent evaluation.
MiniMax, a Shanghai-based AI company that listed on the Hong Kong Stock Exchange on January 9, 2026. CNBC reported the debut as "MiniMax doubles in Hong Kong debut, marking yet another Chinese AI listing." MiniMax also builds the Hailuo video models and the MiniMax Agent product.
Yes, as of September 6, 2026. MiniMax's API release notes list no M-series language model after M3 (June 1, 2026), and the models guide still puts M3 first among current models. Everything MiniMax has shipped since is generative media rather than a successor language model: a new video model on July 31, 2026 and MiniMax Music 3.0 on August 13, 2026.
Not by MiniMax, as of September 6, 2026. MiniMax's models guide classifies MiniMax-M2.5 and MiniMax-M2.5-highspeed as legacy, but both remain callable by exact model ID and still carry a rate card on the pay-as-you-go pricing page, and MiniMax has published no retirement date for them. The deprecations that exist are host-level: Fireworks removed MiniMax M2.5 from serverless on June 17, 2026 with the instruction "Migrate to MiniMax M2.7." If you are planning a migration off M2.5, your host's schedule binds before MiniMax's does.

Explore More Models

Discover other AI models and compare their capabilities.