MiniMax M3 Release Date — and M3.1 Flash Preview

Quick answer. MiniMax M3 was released on 1 June 2026, with open weights on Hugging Face from 2 June 2026. As of 5 October 2026 a successor does exist: MiniMax-M3.1-Flash-Preview is live in MiniMax's API, but subscription-only and with no open weights. There is still no M3 Pro. M3 API pricing is $0.30 / $1.20 per million tokens.

Re-verified against MiniMax's model card, Anthropic-compatible API reference, M Plan documentation, Hugging Face org listing and Artificial Analysis on 5 October 2026. Every date, price and score below comes from a primary source checked on that date.

If you searched for a MiniMax M3 release date, you almost certainly also want to know whether something newer has landed since. That answer changed recently, and quietly: MiniMax has put an M3.1 Flash Preview into its API without a launch post, without open weights and without a pay-as-you-go price. Both answers are below, along with the full M-series timeline and everything MiniMax has shipped in the four months since M3.

FieldStatus as of 5 October 2026
ModelMiniMax-M3
Announced1 June 2026 (evening of 31 May, US Eastern)
Open weights published2 June 2026, MiniMaxAI/MiniMax-M3 on Hugging Face
Current open-weights versionM3 — no point release; repo last updated 23 July 2026
M3.1Exists as MiniMax-M3.1-Flash-Preview — API-only, M Plan / MiniMax Code subscriptions only, no open weights, no pay-as-you-go price
M3 ProDoes not exist. Not on Hugging Face, not in the API model list
Architecture428B total / ~23B active Mixture-of-Experts, natively multimodal
Context window1,000,000 tokens (max output 262,144 tokens)
Licenceminimax-community — open weights, MiniMax's own licence (not MIT/Apache)
API price$0.30 in / $1.20 out per 1M tokens (≤512K prompt); $0.60 / $2.40 above 512K
Newest MiniMax language modelMiniMax-M3.1-Flash-Preview — in the API model list, with no dated announcement. Newest open weights are MiniMax Music 3, 7 Aug 2026

When was MiniMax M3 released?

MiniMax announced M3 on 1 June 2026. The launch went live on the evening of Sunday 31 May US Eastern time, which is why some coverage dates it to 31 May — it is the same event in two timezones. VentureBeat's launch piece is dated 1 June 2026.

The open weights followed almost immediately. The MiniMaxAI/MiniMax-M3 repository was created on 2 June 2026 at 07:49 UTC, with the quantised MiniMax-M3-MXFP8 repo created one minute later. That is the single most reliable release-date artefact available, because Hugging Face timestamps repository creation and does not let it be backdated.

The technical paper behind the model — MiniMax Sparse Attention, arXiv:2606.13392 — was submitted later, on 11 June 2026. If you are trying to cite M3 academically, that is the reference; if you are trying to date the release, use 1–2 June.

Here is where M3 sits in the full MiniMax lineage. Dates are Hugging Face repository creation dates from the official MiniMaxAI org, checked today.

ModelWeights publishedWhat it is
MiniMax-Text-01 / VL-0112 Jan 2025First open-weights generation
MiniMax-M1 (40k / 80k)5–13 Jun 2025First reasoning-focused M-series model
MiniMax-M222 Oct 2025Agentic turn; 200K context, function calling
MiniMax-M2.120 Dec 2025230B total / 10B active; stronger reasoning
MiniMax-M2.512 Feb 2026Code generation and refactoring focus
MiniMax-M2.79 Apr 2026Last M2-line release; still on the API today
MiniMax-M32 Jun 2026Current flagship. 428B/23B MoE, 1M context, multimodal
MiniMax-M3-MXFP82 Jun 2026Official MXFP8 quantisation of M3 for self-hosting
MiniMax-H328 Jul 2026Video generation model — a different product line, not an M3 successor
MiniMax-Music37 Aug 2026Music generation model
MiniMax-M3.1-Flash-PreviewNo weights publishedCurrent flagship text model on M Plan subscriptions. API-only; appeared in the docs with no release-notes entry

Is there a MiniMax M3 Pro or an M3.1?

M3.1 now exists. M3 Pro still does not. Both halves need unpacking, because M3.1 arrived in a shape most release trackers missed entirely.

MiniMax-M3.1-Flash-Preview sits at the top of the Language table on MiniMax's models overview and is a fully documented model ID in the Anthropic-compatible API reference. MiniMax describes it as a "frontier multimodal coding model with 1M context window and tunable thinking depth". What is unusual is everything around it:

  • Subscription-only. The docs carry a single note: "MiniMax-M3.1-Flash-Preview is available only through M Plan and MiniMax Code for now." It has no row on the pay-as-you-go rate card, so you cannot buy it by the token.
  • No open weights. The MiniMaxAI org has published nothing since MiniMax-Music3 on 7 August 2026. There is no M3.1 repository, and no MiniMax-M3-Pro either.
  • No announcement. MiniMax's own model release notes still end at MiniMax H3 on 31 July 2026 — no blog post, no dated entry, no benchmark table.
  • No third-party listing. It is absent from OpenRouter's model catalogue and has no Artificial Analysis page, so there is no independent score for it yet, only MiniMax's own description.

Three documented API behaviours matter before you point a harness at it:

  • It always thinks. Sending thinking: {"type": "disabled"} returns a 400 with error 2013 — "requires adaptive thinking". M3, by contrast, has thinking off by default and opts in with adaptive.
  • Thinking depth is a dial. output_config.effort accepts low, medium, high, xhigh and max, defaulting to max when omitted; none is rejected. That is a real cost lever, because higher effort means more thinking tokens and higher latency.
  • Same multimodal surface as M3. Text, image and video input plus tool use and thinking blocks; images up to 10 MB, videos up to 50 MB inline or 512 MB through the Files API.

There is a second, circumstantial signal that M3.1 Flash Preview has been taking real traffic. Between 23 September and 5 October 2026 an unattributed model called space-bunny-alpha was served free on OpenRouter under the provider name "Stealth", with a 1M-token context window and video input — a parameter surface that lines up closely with M3.1 Flash Preview. We laid out that evidence in our space-bunny-alpha breakdown. No lab has claimed it, so treat the attribution as a strong inference, not a fact.

M3 Pro, meanwhile, genuinely does not exist — not on Hugging Face, not in the API model list, not on MiniMax's product site. The only official M3 variant is MiniMax-M3-MXFP8, the officially quantised build published the same day as the base model. It is the same model at lower precision for self-hosting, not a stronger tier, and MiniMax's own deployment recipe uses it.

The practical consequence: if you want M3.1, you buy an M Plan subscription. If you want weights you can run yourself, M3 is still the newest thing MiniMax has released.

What has MiniMax shipped since M3?

One new language model — M3.1 Flash Preview, covered above — plus a lot of work in other modalities. The M3 repository itself was last updated on 23 July 2026 (documentation and config, not a version bump). Every open-weights release since June has been non-text:

  • MiniMax H3 (28 July 2026) — an omni-modal video generation model. It takes text, image, video and audio context and generates video with synchronised stereo audio at up to 2K resolution, 24 FPS, in 4–15 second clips, with stable support for 11 languages. Architecturally it is nothing like M3: a dense 33B-parameter transformer (the H3-Omni-Transformer), not a sparse MoE. It ships under a separate MiniMax H3 Community License Agreement and has already pulled over 9.6 million all-time downloads. An H3 Max high-speed variant is served through fal.ai. H3 is not "M3 for reasoning, but newer" — if you are choosing a text or coding model, H3 is irrelevant to you.
  • MiniMax Music 3 (7 August 2026) — music generation, exposed as music-3.0. Note that MiniMax's docs state the paid music APIs are unavailable to new users after 20 August 2026.
  • Speech 2.8 — the current text-to-speech line, with hd and turbo tiers covering 40 languages.

So the honest summary: MiniMax spent the summer pushing into video and audio, and its one post-M3 language model arrived as a gated preview rather than an open-weights launch. If you track MiniMax by Hugging Face releases alone, you missed it.

What is MiniMax M3, technically?

Per the official model card, M3 is a 428B-parameter Mixture-of-Experts model that activates roughly 23B parameters per token, trained natively multimodal — mixed text, image and video from the first training step rather than a vision adapter bolted on afterwards. It serves a 1,000,000-token context window and can emit up to 262,144 completion tokens in a single response.

The engineering that makes the long context affordable is MiniMax Sparse Attention (MSA). The MSA paper describes it as blockwise sparse attention built on top of Grouped Query Attention, split into an Index Branch and a Main Branch. The reported effect at 1M context is a 28.4× reduction in attention computation, translating to 14.2× prefill and 7.6× decoding wall-clock speedups on H800 hardware in the paper's own measurements. MiniMax's model card quotes the end-to-end figure differently — 9× prefill and 15× decode speedup versus M2 at 1M context, with per-token compute cut to roughly 1/20.

Two behavioural details matter if you are actually building on it:

  • Three reasoning modes: enabled, adaptive, and disabled. Adaptive lets the model decide when a query warrants a thinking pass, which is the practical default for mixed workloads.
  • Interleaved thinking across tool calls. M3 reasons between each round of tool interaction rather than only at the start. MiniMax's tool-use documentation is explicit that you must append the complete assistant message — including the thinking content — back into conversation history, or the reasoning chain breaks and quality degrades. This is the most common integration mistake with M3.

Recommended sampling parameters from the model card are temperature=1.0 and top_p=0.95. Supported inference stacks are SGLang, vLLM, Transformers, KTransformers, unsloth and ATOM.

How does MiniMax M3 actually benchmark?

Separate these two tables. The first is what MiniMax reports; the second is what an independent evaluator measures. They tell noticeably different stories, and the gap is the most important thing on this page.

Vendor-reported, from the M3 model card and the 1 June launch coverage:

BenchmarkMiniMax M3 score
SWE-bench Verified80.5%
SWE-bench Pro59.0
MMMU Pro78.1%
Terminal-Bench 2.166.0
MCP Atlas74.2
BrowseComp83.5
Skillsbench V1.153
Apex Agents27.7
LHTB38.5 mean reward (3 of 46 tasks solved)

At launch that 59.0 on SWE-bench Pro was the headline, because it put M3 ahead of GPT-5.5 and Gemini 3.1 Pro on that specific benchmark at a fraction of their token price. That framing was accurate on 1 June.

Independently measured, from Artificial Analysis — Intelligence Index v4.3.2, a composite of independent reasoning, coding and knowledge evaluations, checked 5 October 2026. One caveat that invalidates most comparison tables you will find elsewhere: AA rebased the index from v4.1.1 to v4.3.2, which moved every model's score because the benchmark basket changed, not because the models did. v4.1.1 and v4.3.2 figures are not comparable and there is no conversion factor. Where a model exposes effort tiers the variant is named below, because the spread inside a single model can exceed the gap between models:

Open-weights model (variant)Intelligence Index v4.3.2Active paramsUSD per Index task
GLM-5.3 (Max)44.7840B$2.01
Kimi K3 (Max)43.59104B$2.00
GLM 5.3 Flash41.8118B$0.25
Qwen3.8 2.4T A95B39.8995B$2.16
DeepSeek V4 Pro 081336.0049B$0.67
Qwen3.8 27B (Xhigh)33.7027B$1.01
Motif 333.5713.2B$5.41
MiniMax-M329.2223B$0.51
Inkling (Xhigh)24.9841B$0.80

M3 scores 29.22 on Intelligence Index v4.3.2, eighth of the nine open-weights models above. We have dropped the leaderboard rank this page used to quote, because AA's field has grown from 212 to 224 models and a rank captured against the old denominator misleads even when the score behind it is right. Four months after launch M3 is not a leader on aggregate intelligence — GLM-5.3, Kimi K3, Qwen3.8 and DeepSeek V4 Pro all score higher. What it still has is cheap tokens and a very large window: $0.51 per Index task, 97 output tokens/sec, a 2.03-second median time to first chunk, and a 1M-token context nothing else on that list matches. Running the entire v4.3.2 Index on M3 costs AA $537.98 — a useful absolute figure, but not comparable across models, because AA runs a different number of tasks per model. That is why the per-task column is the one to compare.

On cost per Index task M3 is second-cheapest in this group, not cheapest. GLM 5.3 Flash is both cheaper ($0.25) and substantially smarter (41.81), which is the strongest single argument against standardising on M3 unless you specifically need the context window.

The lesson is a general one for open-weights models: a launch-day benchmark table has a shelf life of roughly one quarter. Re-check before you standardise on anything.

How much does MiniMax M3 cost?

Current pay-as-you-go pricing from MiniMax's official rate card, verified today. Note that the price now tiers on prompt length — a change from launch, when the $0.30 / $1.20 rate was announced as a one-week promotion against a $0.60 / $2.40 list price. The cheap rate stuck: MiniMax's rate card now labels it "Permanent 50% off", with the $0.60 / $2.40 figures struck through.

TierInput / 1MOutput / 1MCache read / 1M
M3 standard, ≤512K prompt$0.30$1.20$0.06
M3 standard, >512K prompt$0.60$2.40$0.12
M3 priority, ≤512K prompt$0.45$1.80$0.09
M3 priority, >512K prompt$0.90$3.60$0.18
MiniMax-M2.7$0.30$1.20$0.06
MiniMax-M2.7-highspeed$0.60$2.40$0.06

The 512K threshold is a real budgeting concern if you are using the 1M window as a selling point: crossing it doubles your rate on the entire request, not just the overflow. For long-document workloads, aggressive prompt caching at $0.06/M is where the savings actually live.

If you would rather pay a flat rate, note that MiniMax has replaced the old Token Plan with M Plan, sold in three tiers: Go at $22/month, Explore at $55/month and Build at $132/month. Tiers are now rated by usage allowance rather than concurrent agents — Explore is 3× Go, Build is 7.5× Go — and annual billing charges ten months for twelve. The detail that matters most on this page: the text model on all three M Plan tiers is M3.1 Flash Preview, not M3. Explore and Build add the H3 video model; Go includes no video model. A subscription key is not interchangeable with a pay-as-you-go API key, and MiniMax's own guidance is that M Plan suits interactive developer use while production traffic belongs on pay-as-you-go.

How do you access or self-host MiniMax M3?

Three routes, in ascending order of effort:

  • Hosted API. MiniMax exposes both an OpenAI-compatible and an Anthropic-compatible endpoint, so most existing clients work with a base-URL swap. There is a documented Claude Code integration path if that is your harness. M3 is also routed through third-party aggregators such as OpenRouter.
  • MiniMax Agent. The vendor's own hosted agent product, if you want to try the model's tool-use behaviour without writing an integration.
  • Self-hosted. This is a serious hardware commitment. MiniMax's official deployment recipe baselines on 8 × NVIDIA B200, single node, tensor-parallel 8, with weights of ~444 GB at the pinned MXFP8 revision (the BF16 build on H200 is ~854 GB). SGLang is the supported framework, via a SHA-pinned Docker image. B300, GB300, GB200, H200 and AMD MI355X / MI350X / MI325X / MI300X configurations are also validated, all single-node.

MiniMax's docs include a warning worth repeating, because it trips people up constantly: "'23B activated parameters' does not mean that the deployment only needs memory for 23B parameters." The full 428B of weights must be resident or sharded. Sparse activation buys you compute, not VRAM.

For quantised community builds, memory maths and consumer-hardware options, see our walkthrough on running MiniMax M3 locally. For API wiring, agent harness setup and an evaluation playbook, the MiniMax M3 developer guide covers the integration path end to end.

How does MiniMax M3 compare to DeepSeek V4, GLM, Qwen and Kimi?

Using the independent numbers above rather than launch-day marketing, the open-weights picture in early October 2026 looks like this:

  • vs DeepSeek V4 Pro 0813 — DeepSeek leads on aggregate intelligence, 36.00 against M3's 29.22 on Intelligence Index v4.3.2, and costs more per Index task ($0.67 vs $0.51). Pick DeepSeek for raw quality per task; pick M3 when volume economics or context length dominate.
  • vs GLM-5.3 — still the strongest all-round case against M3. GLM-5.3 (Max) scores 44.78 at $2.01 per Index task, and GLM 5.3 Flash scores 41.81 at $0.25 — cheaper and substantially smarter than M3 on v4.3.2. If your workload fits inside a normal context window, Flash is hard to argue with. We put the two coding lines head to head in GLM vs MiniMax M3 for coding.
  • vs Qwen3.8 — the 2.4T A95B build scores 39.89 on v4.3.2, and even the 27B at its Xhigh effort tier scores 33.70. Both beat M3. Qwen wins on quality and on ecosystem breadth; M3 wins on price and on context.
  • vs Kimi K3 — Kimi K3 (Max) scores 43.59, near the top of the open-weights field, but at $2.00 per Index task it costs roughly four times what M3 costs for the same benchmark work. Different budget bracket entirely.

The defensible claim for M3 today is narrow but real: it is among the cheapest credible open-weights agentic models, and it is the only one in that group with a million-token window. If your bottleneck is fitting an entire codebase or a long document set into one prompt, nothing on that list substitutes for it. If your bottleneck is answer quality per task, several models now beat it. For a broader cross-model view, see our DeepSeek V4 vs Qwen, Kimi and MiniMax comparison.

Companion guide

For where M3 sits among open-weight models — alternatives, licensing trade-offs, and how to choose what to self-host in 2026 — see our open-source LLMs landscape for 2026.

Should you build on MiniMax M3 today?

A decision rule rather than a recommendation:

  • You need more than 256K of context. Use M3. Nothing else in the open-weights tier offers 1M tokens, and the MSA architecture is what makes it affordable rather than theoretical. Just budget for the 512K price cliff.
  • You are optimising cost per token at high volume. M3 at $0.30 / $1.20 per million is near the floor for its intelligence band — but benchmark GLM 5.3 Flash first. Flash costs $0.25 per Index task against M3's $0.51 and scores 41.81 against M3's 29.22. Flash caps out well below 1M context, which is the only reason to prefer M3 on cost grounds.
  • You are optimising answer quality. M3 is not the pick in October 2026. GLM-5.3, Kimi K3, Qwen3.8 and DeepSeek V4 Pro all score higher on independent aggregate evaluation.
  • You are still on M2.7. No urgency. M2.7 remains on the API at identical pricing and is not deprecated. Migrate when you need the context window or the multimodal input, not on principle.
  • You want M3.1. It exists, but only inside an M Plan subscription or MiniMax Code. There is no pay-as-you-go price, no published weights and no independent benchmark, so you cannot put it behind production traffic, self-host it, or compare it on paper. Evaluate it; do not build a roadmap on it yet.

Whatever you pick, keep the inference layer model-agnostic. The four months between M3's launch and today produced four open-weights models that outscore it, plus a successor you can only rent by the month. That cadence is not slowing down, and the teams that handle it well are the ones for whom a model swap is a config change rather than a refactor. If you need that kind of evaluation and self-hosting experience on your team, Codersera matches you with vetted remote developers who have shipped it in production.

FAQ

When was MiniMax M3 released?

MiniMax M3 was announced on 1 June 2026, on the evening of 31 May US Eastern time. The open weights were published to the MiniMaxAI/MiniMax-M3 Hugging Face repository on 2 June 2026 at 07:49 UTC, alongside the quantised MiniMax-M3-MXFP8 build. The MiniMax Sparse Attention paper describing the architecture followed on 11 June 2026.

Is there a MiniMax M3 Pro?

No. As of 5 October 2026 there is no MiniMax M3 Pro on Hugging Face, in the MiniMax API model list, or on the MiniMax product site. The only official M3 variant is MiniMax-M3-MXFP8, the same model quantised to MXFP8 precision for self-hosting — not a higher-capability tier. The model that does sit above M3 is MiniMax-M3.1-Flash-Preview, and it is subscription-only with no published weights. Anyone quoting M3 Pro specifications is describing something that has not been released.

Is MiniMax M3.1 out?

Partly. MiniMax-M3.1-Flash-Preview is live in MiniMax's API model list and Anthropic-compatible API reference as of 5 October 2026, described as a frontier multimodal coding model with a 1M-token context window and tunable thinking depth. But it is reachable only through an M Plan subscription or MiniMax Code: there is no pay-as-you-go price, no Hugging Face repository, no release-notes entry and no independent benchmark score. A full, open-weights M3.1 has not shipped.

Can you self-host MiniMax M3.1 Flash Preview?

No. MiniMax has published no weights for M3.1 Flash Preview, so there is nothing to download. The newest MiniMax language model you can self-host is M3, via the MiniMaxAI/MiniMax-M3 or MiniMax-M3-MXFP8 repositories. M3.1 Flash Preview is reachable only through MiniMax's hosted API with an M Plan subscription key, which is also why no third party has benchmarked it yet.

Is MiniMax M3 open source?

M3 is open weights, not open source in the OSI sense. The full 428B weights are downloadable from Hugging Face, but they ship under MiniMax's own minimax-community licence rather than MIT or Apache 2.0. Read the licence file before commercial deployment. Training data and training code are not published.

How much does MiniMax M3 cost?

API pricing is $0.30 per million input tokens and $1.20 per million output tokens for prompts up to 512K, doubling to $0.60 / $2.40 above 512K. Cached reads are $0.06 per million. A priority tier costs 1.5× standard. MiniMax's M Plan subscriptions run $22/month (Go), $55/month (Explore) and $132/month (Build), but their text model is M3.1 Flash Preview rather than M3.

How does MiniMax M3 compare to DeepSeek V4?

DeepSeek V4 Pro 0813 scores 36.00 on Artificial Analysis's Intelligence Index v4.3.2 against M3's 29.22, so DeepSeek is meaningfully stronger on aggregate reasoning and coding quality. M3 is cheaper per benchmark task ($0.51 versus $0.67) and far cheaper per token ($0.30 / $1.20 per million against DeepSeek's $1.32 / $3.96), and it offers a 1M-token context window. Choose DeepSeek for quality, M3 for cost and context length.

What hardware do you need to self-host MiniMax M3?

MiniMax's reference deployment is 8 × NVIDIA B200 in a single node at tensor-parallel 8, using SGLang and the MXFP8 weights, which occupy roughly 444 GB. The BF16 build on H200 is about 854 GB. B300, GB300, GB200 and AMD MI355X / MI350X / MI325X / MI300X are also validated. The 23B active-parameter figure describes compute, not memory — all 428B parameters must be resident.

Is MiniMax H3 a newer version of M3?

No. H3, published on 28 July 2026, is a video generation model — a dense 33B-parameter omni-modal transformer that produces video with synchronised stereo audio at up to 2K resolution in 4–15 second clips. It is a separate product line with its own licence. If you are choosing a model for text, coding or agent work, H3 is not a candidate; M3 remains MiniMax's newest open-weights language model, with M3.1 Flash Preview available only through a subscription.