GLM-5.2 vs MiniMax M3: Open-Weights Coding (2026)
Quick answer. On the one board that scores both, GLM-5.2 leads MiniMax M3 by 4.5 points — 33.71 against 29.22 on Artificial Analysis's Intelligence Index v4.3.2 — but only at GLM's Max reasoning tier. Run GLM-5.2 with reasoning off and it scores 22.43, below M3. M3 wins on price ($0.30/$1.20 per M tokens vs $1.40/$4.40) and is the only one of the two that accepts image and video input. GLM-5.2's weights are MIT; M3's are under MiniMax's own community licence.
Update — October 5, 2026: GLM-5.3-Flash dissolves the trade-off this article is about. When we wrote this, choosing GLM meant paying 4–5× more per token than M3, and choosing M3 meant giving up coding-leaderboard standing. GLM-5.3-Flash, released 25 August 2026 under MIT, is both stronger and cheaper than M3 on every figure here: 41.81 on Artificial Analysis's Intelligence Index v4.3.2 against M3's 29.22, at $0.15/$0.50 per million tokens against M3's $0.30/$1.20, and $0.25 per Index task against M3's $0.51. M3's surviving advantages are real but narrower than they were: native image and video input, which no GLM text model has, and MSA's long-context throughput.
Two corrections to the box that used to sit here. GLM-5.3 does have public weights — they landed 25 August 2026 and have 1.4M downloads in the last 30 days — and a metered API went live 18 August, so Artificial Analysis has scored it: 44.78 at the Max tier, 34.30 at the default tier. But GLM-5.3 proper is not MIT; it ships under Z.ai's own glm-5.3 licence. Only GLM-5.3-Flash is MIT. See the GLM-5.3 launch guide.
GLM-5.2 vs MiniMax M3: at a glance
GLM-5.2 (from Z.ai) and MiniMax M3 (from Shanghai lab MiniMax) are two of the most-discussed open-weight releases of mid-2026. Both target the same buyer — teams that want strong coding without paying closed-frontier API rates — but they got there with very different design choices. Here is the shape of the matchup before we go deep.
| GLM-5.2 | MiniMax M3 | |
|---|---|---|
| Maker / release | Z.ai — 13 June 2026 | MiniMax — 1 June 2026 |
| Architecture | MoE, ~744B total / ~40B active | MoE, ~428B total / ~23B active |
| Attention | Dense-style long-context (1M) | MiniMax Sparse Attention (MSA) |
| Context window | 1,000,000 tokens (131K output) | 1,000,000 tokens |
| Modality | Text only | Text + image + video input |
| License | MIT (zai-org/GLM-5.2 on HF) | MiniMax community licence — open weights, but not a standard OSI licence (MiniMaxAI/MiniMax-M3) |
| API price (in / out, 1M) | $1.40 / $4.40 ($0.26 cached in) | $0.30 / $1.20 — MiniMax now labels this a permanent 50% discount, not a launch promo. Doubles to $0.60 / $2.40 above 512K input tokens. |
| Intelligence Index v4.3.2 | 33.71 (Max tier) / 22.43 (reasoning off) | 29.22 |
| Cost per Index task | $1.47 | $0.51 |
| Headline strength | Stronger text coding at its Max reasoning tier; MIT weights | Multimodality + efficiency + price |
How do the coding benchmarks actually compare?
Start with the caveat that matters most: the two labs have not run on a single shared harness, and they even report different Terminal-Bench versions. So treat any "M3 beats GLM by X points" claim — including ones you see on Twitter — with suspicion. What we can do honestly is line up each model's own published numbers, keep the harnesses separate, and read the direction of travel.
| Benchmark (harness) | GLM-5.2 | MiniMax M3 |
|---|---|---|
| Terminal-Bench (GLM's reported run, per Cline) | >80% (first open model past 80%) | not reported on this version |
| Terminal-Bench 2.1 (MiniMax's reported run) | not reported on this version | 66.0% |
| SWE-Bench Pro | not separately published | 59.0% |
| BrowseComp (agentic browsing) | not published | 83.5% |
| SVG-Bench | not published | 63.7% |
| Artificial Analysis Intelligence Index v4.3.2 (Oct 2026) | 33.71 (Max tier) · 22.43 reasoning off | 29.22 |
| Cost per Index task (Artificial Analysis) | $1.47 | $0.51 |
| Median output speed (Artificial Analysis) | 94.35 tok/s at Max tier | 97.22 tok/s |
| Design Arena (web / UI, Elo) | Elo 1360, top open-weights entry as of mid-2026 | listed |
Read it carefully: the GLM and MiniMax Terminal-Bench figures are on different versions of the harness and are not comparable as a single row, which is why they sit on separate lines above. The Artificial Analysis rows are comparable, and they are the only ones on this page that are — AA runs its own harness against both models. Correcting an earlier version of this article: M3 is on that board. It scores 29.22 on Intelligence Index v4.3.2 against GLM-5.2's 33.71, so GLM leads by 4.5 points.
The qualifier matters more than the gap. 33.71 is GLM-5.2's Max reasoning tier. With reasoning off, AA measures GLM-5.2 at 22.43 — nearly 7 points below M3. AA publishes a single figure for M3 and multiple tiers for GLM, so any “GLM beats M3” claim is only true if you are actually paying for Max-tier reasoning, which costs more tokens and drops throughput. If your agent runs GLM-5.2 in its cheap non-reasoning configuration to control spend, M3 is the stronger model and also the cheaper one. Note too that AA rebased this index from v4.1.1 to v4.3.2 during 2026 with no conversion factor, so these figures cannot be compared against scores quoted in older write-ups of either model.
MiniMax M3 publishes its own respectable coding scores (59.0% SWE-Bench Pro, 66.0% Terminal-Bench 2.1) and a notably strong agentic-browsing result (83.5% BrowseComp). GLM-5.2 doesn't publish SWE-Bench Pro or BrowseComp, so those aren't head-to-head either — they're simply areas where M3 has a number on the board and GLM doesn't. The practical takeaway: for terminal-style and UI coding measured on the boards GLM appears on, GLM-5.2 is the front-runner; for multimodal and agentic-browsing tasks, M3 is the one with the published evidence.
Is the 1M-token context window actually useful?
Both models advertise a 1,000,000-token context window, which on paper means you can fit a large codebase into a single request. In practice the two get there differently, and that difference shows up on your bill and your latency graph.
GLM-5.2 serves 1M context in its glm-5.2[1m] variant with up to 131,072 output tokens per response — a 5× jump over GLM-5.1's 200K window. It is a strong long-context model, but the KV cache at 1M is large, so most long-context use runs through Z.ai's API or a provisioned cluster rather than a single workstation.
MiniMax M3's headline architectural claim is MiniMax Sparse Attention (MSA): a sparse-attention scheme built on a Grouped-Query Attention backbone that does block-level selection over real, uncompressed key-values. MiniMax reports roughly 9.7× faster prefill and 15.6× faster decoding at 1M-token context versus the previous generation, with per-token compute around one-twentieth of MiniMax M2. If your workload routinely fills the context window — long agent traces, whole-repo refactors, multi-document analysis — M3's long-context economics are meaningfully better on the numbers MiniMax has published.
What do the token economics look like?
This is where MiniMax M3 lands its biggest advantage — and the pricing picture has firmed up since launch. M3 is $0.30 / 1M input and $1.20 / 1M output, and MiniMax's own pricing page now labels that a “permanent 50% off” rather than a launch promotion, so the $0.60 / $2.40 list price is best read as a reference figure rather than a cliff you are about to fall off. There is a real cliff elsewhere: above 512K input tokens M3 switches to $0.60 / $2.40 on the standard tier, exactly double. Since M3's long-context economics are its own headline claim, that threshold is worth knowing before you plan a workload around the full 1M window. GLM-5.2's metered API runs $1.40 / 1M input and $4.40 / 1M output, with cached input at $0.26 / 1M, flat across the whole window.
At list prices, M3 is roughly 4–5× cheaper per output token than GLM-5.2. For high-volume agent workloads where output tokens dominate, that gap compounds quickly. GLM-5.2's counterweight on cost is its $0.26 cached-input rate: agent loops that re-send the same system prompt and codebase context on every turn can recover a large fraction of the difference. Both also undercut the closed-frontier APIs by a wide margin on published per-token pricing; our GLM-5.2 complete guide breaks down the cost math in detail.
The honest read: if raw cost-per-token is your binding constraint, M3 beats GLM-5.2 — and Artificial Analysis's cost-per-Index-task figures agree, at $0.51 for M3 against $1.47 for GLM-5.2 at Max. If your agent is cache-heavy, GLM-5.2's $0.26 cached-input rate narrows that. But the 2026 answer to “cheapest capable open-weights coder” is neither of them: GLM-5.3-Flash is $0.15 / $0.50 with a 41.81 index score and $0.25 per Index task, which is cheaper than M3 and 12.6 points stronger.
Self-hosting: the two paths compared
Both models are open-weight, so you can run either on your own hardware — but the bill of materials differs because the models are different sizes.
- GLM-5.2 (~744B / 40B active). Unsloth's dynamic GGUF quants make it tractable: per Unsloth's dynamic-quant benchmarks, a 2-bit dynamic quant retains most of its full-precision quality after shrinking the weights to about 238 GB, which fits a 256 GB unified-memory Mac or a single 24 GB GPU with CPU offload. Full-precision serving wants an 8×H100-class cluster. Day-0 support landed in SGLang, vLLM, and llama.cpp.
- MiniMax M3 (~428B / 23B active). Fewer total and active parameters means a lighter footprint and faster decode at the same quant level; Unsloth shipped local quants shortly after the weights landed, and the model runs through llama.cpp / Unsloth Studio. The 23B active-parameter count keeps decode fast on modest hardware.
If you want the model topping the open-weights coding boards and you have the memory, self-host GLM-5.2. If you want the lighter, faster, multimodal model that's kinder to a single workstation, M3 is the easier self-host.
Does multi-modal matter here?
For a lot of real coding work in 2026, yes — and this is M3's structural advantage. MiniMax M3 accepts text, image, and video input natively; GLM-5.2 is text-only. If your workflow includes screenshot-to-code, debugging from a screen recording, reading design mockups, or building UI from a Figma export, M3 can see the input and GLM-5.2 cannot. With GLM-5.2 you would otherwise have to bolt a separate vision model onto the pipeline.
If your coding is pure text — repos, terminals, logs, specs — GLM-5.2's lack of vision costs you nothing, and you get the higher coding-leaderboard standing in return.
Who should pick GLM-5.2?
- You want the open-weights model that currently tops the coding leaderboards it appears on, for text work.
- Your work is text-first: repositories, CLIs, refactors, test generation, spec-to-code.
- You run cache-heavy agent loops where the $0.26 cached-input rate offsets the higher sticker price.
- You're building web or UI and want the current Design Arena #1.
- You have (or rent) the memory to self-host a ~744B MoE and want MIT-licensed weights you control.
Who should pick MiniMax M3?
- Your workflow is multimodal — screenshots, design mockups, video, or mixed image+code tasks.
- Cost-per-token is the binding constraint and you ship high output volume.
- You fill the 1M context routinely and want MSA's faster, cheaper long-context inference.
- You want a lighter self-host (fewer active parameters) on a single workstation.
- You lean on agentic browsing, where M3's published 83.5% BrowseComp is strong.
The decision in a line
Pick GLM-5.2 if you need MIT-licensed weights pinned on your own hardware and you will actually run it at the Max reasoning tier. Pick MiniMax M3 if you need image or video input, or if GLM-5.2's non-reasoning tier is all your budget allows — because at that tier M3 is both stronger and cheaper.
And a third option that did not exist when this comparison was written: if the only reason you are weighing M3 is price, GLM-5.3-Flash wins outright — MIT, $0.15 / $0.50, 41.81 on Intelligence Index v4.3.2, $0.25 per Index task. It beats M3 on capability and on cost simultaneously. M3 keeps the multimodal lane to itself.
Architecture details that matter for capacity planning
GLM-5.2 is a Mixture-of-Experts model with approximately 744 billion total parameters and about 40 billion active per token (some vendor write-ups quote 753B total). It exposes a usable 1M-token context in the glm-5.2[1m] variant. The weights are MIT-licensed on Hugging Face under zai-org/GLM-5.2; the training code and full technical report were not published at launch, so "open weights" is more precise than "open source."
MiniMax M3 is a ~428B-total / ~23B-active MoE whose defining feature is MSA — sparse attention over uncompressed KV on a GQA backbone, tuned for cheap 1M-context inference. MiniMax positions M3 as the first open-weight model to combine strong coding, a 1M context window, and native multimodality in one checkpoint.
FAQ
Is GLM-5.2 or MiniMax M3 better for coding?
GLM-5.2 at its Max reasoning tier: 33.71 against M3's 29.22 on Artificial Analysis's Intelligence Index v4.3.2, the one harness that has run both. With reasoning off GLM-5.2 drops to 22.43 and M3 is clearly better, so the answer depends on which GLM tier you are actually paying for. M3 also owns the multimodal lane outright and publishes its own numbers on separate harnesses (59.0% SWE-Bench Pro, 66.0% Terminal-Bench 2.1, 83.5% BrowseComp) that have no GLM-5.2 counterpart. Neither is the strongest open-weights coder available in October 2026 — GLM-5.3-Flash is, at 41.81.
Which is cheaper, GLM-5.2 or MiniMax M3?
MiniMax M3, at $0.30 / $1.20 per 1M tokens against GLM-5.2's $1.40 / $4.40 — roughly 3.7× cheaper per output token. MiniMax describes that rate as a permanent 50% discount off a $0.60 / $2.40 list, so it is not a promo about to lapse; but note the price doubles to $0.60 / $2.40 once a request exceeds 512K input tokens. GLM-5.2's $0.26 cached-input rate narrows the gap for cache-heavy agent loops. On Artificial Analysis's cost-per-Index-task measure, which accounts for how many tokens each model actually burns to finish the work, M3 is $0.51 against GLM-5.2's $1.47.
Do both models support a 1M-token context window?
Yes. Both advertise 1,000,000-token context. MiniMax M3 uses its MSA sparse attention for faster, cheaper long-context inference; GLM-5.2 serves 1M context in its glm-5.2[1m] variant with up to 131K output tokens.
Does GLM-5.2 support image or video input?
No. GLM-5.2 is text-only. MiniMax M3 natively accepts text, image, and video, which is the main reason to choose M3 for screenshot-to-code or design-to-code work.
Can I run both models locally?
Yes. Both publish weights. GLM-5.2 (~744B) runs via Unsloth dynamic GGUF — a 2-bit quant near 238 GB fits a 256 GB Mac or a 24 GB GPU with offload. MiniMax M3 (~428B / 23B active) is a lighter self-host thanks to fewer active parameters. The licences differ and it matters for commercial deployment: GLM-5.2 is plain MIT, while M3 ships under MiniMax's own community licence — read it rather than assuming Apache or MIT terms.
Are these benchmark comparisons apples-to-apples?
No. The labs report different Terminal-Bench versions and haven't run on a shared harness, so the tables above keep each model's published numbers separate rather than asserting a single-harness head-to-head. Treat any precise "X beats Y by N points" coding claim with caution.
Related reading
- GLM-5.2 complete guide (2026) — architecture, benchmarks, and the three ways to run it.
- MiniMax M3: developer guide to the open-weight 1M-context frontier.
- DeepSeek V4 complete guide (2026) — the other MIT-licensed MoE giant.
- Open-source LLMs landscape (2026) — where GLM-5.2 and M3 sit in the wider field.
Shipping with open-weight models? Codersera helps you extend your engineering team with vetted remote developers experienced in open-weight LLM deployment — GLM-5.2, MiniMax M3, DeepSeek V4 and the rest — from self-hosting and quantisation to agent pipelines. Extend your engineering team →