Kimi K2.6 vs DeepSeek V4 vs GLM-5.1: Coding Verdict
Quick answer. On Artificial Analysis's current Intelligence Index v4.3.2 (read 5 October 2026) DeepSeek V4-Pro 0424 leads this trio at 30.45, with Kimi K2.6 at 26.98 and GLM-5.1 at 26.06 — and V4-Pro is also the cheapest per completed task at $0.12 against $0.80 and $1.02. Pick DeepSeek V4-Pro as the default, Kimi K2.6 for cache-heavy agent loops, GLM-5.1 for a clean MIT licence. None of the three is near the top of open weights any more.
Heads-up — a newer generation of this matchup exists. Since this comparison was written, Moonshot shipped Kimi K2.7-Code (June 11) and Z.ai shipped GLM-5.2 (June 13) and then GLM-5.3 (August 14). The analysis below still holds for K2.6/GLM-5.1, but for the current-gen verdict read: Kimi K2.7 vs DeepSeek V4, Kimi K2.7 vs GLM 5.2, and the Kimi K2.7 complete guide, and the GLM-5.3 launch guide.
Updated 2026-05-23: DeepSeek V4-Pro pricing flipped to permanent $0.435 / $0.87 per 1M input/output (cache-hit $0.003625/M) on 2026-05-22, per api-docs.deepseek.com. Table and cost-per-task section updated; rest of analysis (benchmarks, positioning) unchanged.
Updated 2026-06-14: GLM 5.2 has shipped. Z.ai launched GLM 5.2 on June 13, 2026 with a 1M-token context window and MIT-licensed open weights arriving the week after launch. Zhipu has not published benchmarks at launch, so the GLM-5.1 numbers below remain the published baseline for this comparison. For head-to-head on the new model: GLM 5.2 vs DeepSeek V4, GLM 5.2 vs Claude Opus 4.8, GLM 5.2 vs GPT-5.5.
Updated 2026-10-05: every Artificial Analysis figure on this page was re-read against Intelligence Index v4.3.2. AA rebased the index from v4.1.1 (and the v4.0 series before it) during 2026, changing both the benchmark basket and the aggregation. It is a reset, not a drift — there is no conversion factor between versions, so an old score of 54 and a new score of 27 can describe the same model on the same day. The composite ordering of this trio changed as a result: DeepSeek V4-Pro now leads it.
By April 2026 the open-weights coding race stopped being a story about catching closed models and became a story about which open model to standardise on. Three names dominate the shortlist: Moonshot's Kimi K2.6, DeepSeek V4 (shipped as V4-Pro and V4-Flash), and Z.ai's GLM-5.1. All three are downloadable, all three are credible against the closed frontier on coding, and all three pick a different hill to win on.
The honest problem: there is no clean three-way canonical leaderboard. Independent harnesses (Artificial Analysis, LMArena) cover the models unevenly, and the vendors each publish numbers on the benchmarks where they look best. This piece separates neutral, independently-measured data from vendor-reported data, says explicitly where neutral data is missing, and ends with a decision block you can act on. The single sourced 3-column table below is the citable asset — read it first.
Want the full picture? Read our continuously-updated DeepSeek V4 cost breakdown — how V4-Flash ($0.14/$0.28) and V4-Pro compare on price and coding to Claude Opus 5, GPT-5.6, Kimi K3 and Gemini 3.5, with real monthly-bill math.
What are Kimi K2.6, DeepSeek V4 and GLM-5.1?
All three are large Mixture-of-Experts models released within a three-week window in April 2026, all with open weights on Hugging Face.
- Kimi K2.6 (Moonshot AI, released 2026-04-20): 1T total parameters, ~32B active per token, 256K context, configurable thinking/instant modes. Released under a Modified MIT license (permissive, with a visible-attribution clause for very large deployments).
- DeepSeek V4 (DeepSeek, public preview 2026-04-24): two hosted variants from one family. V4-Pro is 1.6T total / ~49B active; V4-Flash is 284B total / ~13B active. Both expose a 1M-token context. Weights and repo are MIT licensed.
- GLM-5.1 (Z.ai, formerly Zhipu, released 2026-04-07): 744B total / ~40B active, 200K context, tuned for long-horizon agentic engineering. MIT licensed.
For background on each model individually, see our Kimi K2.6 complete guide and our DeepSeek V4 complete guide; this article is the head-to-head.
How do they compare on the master benchmark table?
This is the load-bearing table. Every cell is labelled (neutral) if it comes from an independent harness (Artificial Analysis or LMArena) or (vendor) if it is self-reported by the model maker. Where a neutral number does not exist, the cell says so explicitly rather than substituting a vendor number silently.
| Dimension | Kimi K2.6 | DeepSeek V4-Pro (+Flash) | GLM-5.1 |
|---|---|---|---|
| AA Intelligence Index v4.3.2, max effort (neutral, read 2026-10-05) | 26.98 | 30.45 (Pro 0424); 24.17 (Flash 0420) | 26.06 |
| AA cost per index task, v4.3.2 (neutral, read 2026-10-05) | $0.80 | $0.12 (Pro); $0.11 (Flash) | $1.02 |
| AA median output tok/s (neutral, read 2026-10-05) | 59.6 | 107.8 (Pro); 81.2 (Flash) | 58.2 |
| AA GDPval-AA Elo (neutral, v4.1.1-era reading, May 2026 — not re-verified) | 1484 | 1554 (Pro) | 1535 |
| LMArena Code Arena Elo (neutral) | not independently reported | not independently reported | 1530 |
| SWE-bench Verified (vendor) | 80.2% | 80.6% Pro / 79.0% Flash | 77.8% |
| SWE-bench Pro (vendor) | 58.6% (vendor claims #1) | close behind, exact figure not cleanly reported | 58.4% (vendor claims #1) |
| Terminal-Bench 2.0 (vendor) | not cleanly reported | 67.9% (Pro) | strong but exact figure not cleanly reported |
| LiveCodeBench Pass@1 (vendor) | not cleanly reported | 93.5% Pro / 91.6% Flash | not cleanly reported |
| Context window | 256K | 1M (both variants) | 200K |
| List price, $ / 1M in→out | $0.95 → $4.00 (AA-verified) | Pro $0.435 → $0.87 (cache-hit $0.003625/M); Flash $0.14 → $0.28 (vendor list, permanent as of 2026-05-22) | $1.40 → $4.40 (vendor list) |
| Architecture | 1T MoE / ~32B active | Pro 1.6T / ~49B; Flash 284B / ~13B | 744B MoE / ~40B active |
| License | Modified MIT | MIT | MIT |
How to read this honestly. The rows you can trust as apples-to-apples are the neutral ones: Artificial Analysis runs the same harness against every model. On v4.3.2 they now agree with each other, which they did not under the older index — DeepSeek V4-Pro leads the composite (30.45), wins the GDPval-AA agentic benchmark (1554), is the fastest of the three (107.8 tok/s), and is the cheapest per completed task ($0.12). Kimi K2.6's composite lead is gone. GLM-5.1 still holds the only independent Code Arena Elo on record (1530). The vendor SWE-bench / Terminal-Bench / LiveCodeBench rows are useful but not directly comparable across vendors because harness configs, scaffolds, and effort settings differ — treat them as each vendor's best-case, not a ranking.
One scoping caveat on the cost row: AA runs a different number of tasks per model, so whole-suite totals are not comparable between columns. Cost per task is, which is why that is the figure quoted.
What does the neutral data actually say?
Strip out the vendor numbers and the picture narrows to three independent signals:
- Artificial Analysis Intelligence Index v4.3.2 (a composite over GDPval-AA, Terminal-Bench Hard, SciCode, AA-LCR, GPQA Diamond and others), max reasoning effort: DeepSeek V4-Pro 0424 30.45, Kimi K2.6 26.98, GLM-5.1 26.06, DeepSeek V4-Flash 0420 24.17. V4-Pro's lead over the next model is 3.5 points — narrow, but it is a lead, and it is the opposite of what the superseded index showed. Kimi K2.6 and GLM-5.1 are within a point of each other and should be treated as tied.
- Where this trio sits in open weights overall. Nowhere near the top any more. On the same v4.3.2 reading, MiMo-V2.6-Pro scores 46.32, GLM-5.3 44.78, Kimi K3 43.59 and GLM 5.3 Flash 41.81 — all open weights, all released after this comparison was written. Treat this page as a decision aid for the April 2026 generation, not as the open-weights leaderboard.
- GDPval-AA (Artificial Analysis's agentic real-world work benchmark): DeepSeek V4-Pro 1554, GLM-5.1 1535, Kimi K2.6 1484. These Elo figures were read in May 2026, during the v4.1.1 index era, and we have not re-verified them against the current site; they are kept because they are directionally consistent with the rebased composite, not because they are fresh. On the current index V4-Pro leads both the composite and this benchmark, so the two signals no longer disagree.
- LMArena Code Arena: GLM-5.1 has an independently confirmed 1530 Elo. As of this writing there is no comparable independent Code Arena number on record for Kimi K2.6 or DeepSeek V4 — so GLM-5.1 is the only one of the three with a public, independent human-preference coding signal. That is a gap in the neutral data, not a point against the other two.
Net: on the strongest neutral evidence available, DeepSeek V4-Pro is now the pick on the numbers — it leads the composite, leads the agentic-work benchmark, generates fastest, and costs least per completed task. Kimi K2.6 and GLM-5.1 are effectively tied behind it, and GLM-5.1 owns the only independent coding-Elo signal. Under the superseded v4.1.1-era index this section read the other way round, with Kimi first; that reading no longer stands.
Which is cheapest per coding task?
Per-token list price is not cost-per-task — verbosity matters. Artificial Analysis publishes exactly this figure: the measured USD cost to complete one task on its Intelligence Index. On v4.3.2 (read 5 October 2026) it reads $0.11 for V4-Flash 0420, $0.12 for V4-Pro 0424, $0.80 for Kimi K2.6 and $1.02 for GLM-5.1.
That first pair is the surprise, and it reverses the cost argument this section used to make. Per token V4-Flash is roughly 3× cheaper than V4-Pro. Per completed task they are level — because Flash burns the extra tokens back in longer reasoning traces and more attempts. Since V4-Pro also scores six index points higher (30.45 vs 24.17) and ships the same 1M context and MIT licence, there is no longer a cost case for choosing Flash over Pro on the API. Flash's remaining real advantages are self-host footprint (284B against 1.6T) and raw throughput ceiling at low effort.
The per-token picture, for the cases where it still matters (fixed-length summarisation, classification, anything where you control output length tightly):
- DeepSeek V4-Flash ($0.14 in / $0.28 out, vendor list) is the cheapest per token — about 3× below V4-Pro and an order of magnitude below Kimi K2.6 and GLM-5.1. It lands within ~1.6 points of V4-Pro on vendor SWE-bench Verified (79.0% vs 80.6%). But note the paragraph above: on AA's measured cost per index task the two DeepSeek variants are level, so pick Flash for output-bounded jobs and for self-hosting, not as a blanket cost saving.
- Kimi K2.6 ($0.95 in / $4.00 out, AA-verified) is mid-priced with an 83% cache-hit discount on input ($0.16 cached) — strong for cache-heavy agent loops that re-send a stable system prompt and repo map every turn. Without that cache hit it is 6.7× V4-Pro's measured cost per index task ($0.80 against $0.12), which is the gap the cache has to close for K2.6 to make economic sense here.
- DeepSeek V4-Pro ($0.435 in / $0.87 out, vendor list — permanent as of 2026-05-22, per api-docs.deepseek.com) is now dramatically cheaper than it was at launch. With the 75% price cut made permanent and a 90%+ cache-hit discount ($0.003625/M on cached input), V4-Pro is roughly 3× the price of V4-Flash and ~2× cheaper than Kimi K2.6 on output — while still being the GDPval-AA neutral leader. Output verbosity remains the real cost driver, but the post-2026-05-22 economics make V4-Pro a much more comfortable default for long-horizon agentic work than it was a week earlier.
- GLM-5.1 ($1.40 in / $4.40 out, vendor list) is the most expensive of the three on both input and output by a wide margin, and also the most expensive per index task at $1.02. That is offset by a 200K context (cheaper to keep full than a 1M window you actually fill) and the strongest licence story — not by price.
Rule of thumb, on current numbers: if cost dominates your decision, DeepSeek V4-Pro 0424 is the answer, not V4-Flash. It is level with Flash per completed task, six index points ahead, faster, and on the same MIT licence and 1M context. Reach for Flash when the binding constraint is the GPU bill for self-hosting or a hard cap on per-token spend; reach for Kimi K2.6 when a high cache-hit rate is doing the heavy lifting.
How do they compare on self-host and license?
All three are genuinely self-hostable, but the GPU bill and license terms differ.
| Kimi K2.6 | DeepSeek V4-Pro / Flash | GLM-5.1 | |
|---|---|---|---|
| License | Modified MIT (attribution clause at very large scale) | MIT (cleanest) | MIT (cleanest) |
| Inference engines | vLLM, SGLang, KTransformers (Moonshot's own) | vLLM, SGLang | vLLM, SGLang, llama.cpp, Unsloth |
| Practical full-precision hardware | ~8× H200-class; 4× H100 viable with native INT4 (QAT) at reduced context | Pro is the heaviest (1.6T); Flash (284B) is the easiest of all three to self-host | ~8× H100 80GB / ~860GB VRAM for the FP8 checkpoint |
| Self-host sweet spot | Teams wanting top intelligence and willing to pin vLLM/SGLang versions + QAT INT4 | V4-Flash: smallest credible model, 1M context, MIT | Sovereignty/compliance teams that need a clean MIT license and a 200K context |
If "clean license, no asterisks" is a hard requirement (legal, regulated, or resale scenarios), GLM-5.1 and DeepSeek V4 are unencumbered MIT; Kimi K2.6's Modified MIT only adds friction at very large deployment scale, but it is an asterisk a lawyer will flag. If "smallest model I can actually run on my own GPUs" is the constraint, DeepSeek V4-Flash at 284B total is the clear answer and still ships the 1M context window.
Companion guide
For the full landscape — every major open model, how the licenses really differ, and where each one fits in a production stack, see our open-source LLMs landscape for 2026.
Which model should you pick?
The decision block. Match the left column to your real constraint, not to the highest leaderboard number.
- Pick Kimi K2.6 if your agent loop is cache-heavy enough to exploit the 83% input-cache discount, or you specifically want Moonshot's instant/thinking mode switch. It no longer holds the composite lead in this trio — 26.98 against V4-Pro's 30.45 on Intelligence Index v4.3.2 — so this is a workload fit, not a capability pick.
- Pick DeepSeek V4-Flash if self-host footprint or a hard per-token ceiling is the binding constraint. At 284B total it is the smallest credible model here, still ships the 1M context, and is roughly 3× cheaper per token than V4-Pro. What it is not any more is the cost pick for API work: AA measures $0.11 per index task against V4-Pro's $0.12, and Flash scores six points lower.
- Pick DeepSeek V4-Pro if you want the default. On current neutral data it leads this trio on the Intelligence Index v4.3.2 (30.45), leads GDPval-AA (1554), generates fastest (107.8 tok/s), ships a 1M context and MIT weights, and costs least per completed task ($0.12) thanks to the May 2026 permanent price flip to $0.435/$0.87 per 1M. It is the one model here that is not a trade-off.
- Pick GLM-5.1 if you need an unencumbered MIT license, a real independent coding signal (1530 Code Arena Elo), and you do not need more than a 200K context. The sovereignty/compliance pick — you pay for it, at $1.02 per index task against V4-Pro's $0.12.
What is the honest caveat about these numbers?
Be skeptical of any clean ranking — including the headline of this article. Three caveats matter:
- Vendor SWE-bench / Terminal-Bench / LiveCodeBench numbers are not cross-comparable. Each vendor runs its own scaffold, retry budget, and effort setting. Kimi and GLM both claim #1 on SWE-bench Pro within 0.2 points of each other — that gap is inside the noise of differing harnesses. Use vendor numbers to confirm a model is in the frontier band, not to rank within it.
- Neutral data is incomplete. Artificial Analysis covers all three on the Intelligence Index, but LMArena's Code Arena has a public independent Elo only for GLM-5.1. "No independent number" is not "a low number" — it is missing data, and we have marked it as such rather than backfilling with a vendor figure. The GDPval-AA Elo row is a May 2026 reading we have not refreshed, and is labelled as such.
- Index versions get rebased, and the numbers are not convertible. Artificial Analysis moved this composite from the v4.0 series through v4.1.1 to v4.3.2, changing the benchmark basket each time. Kimi K2.6 read 54 on the old basket and 26.98 on the current one without the model changing at all. Any AA score you see quoted without a version stamp is unauditable — including, until this update, the ones on this page. Check the version before you compare two figures.
- Pricing moves and varies by route. DeepSeek V4-Pro pricing was a 75% time-window promo at launch; on 2026-05-22 the company made it permanent at $0.435/$0.87 per 1M (cache-hit $0.003625/M). Kimi K2.6 and GLM-5.1 list prices have held steady. Third-party hosts still re-price all three, and DeepSeek may run further promos on top of the new base — re-check the vendor and Artificial Analysis pages before you commit a budget.
The defensible conclusion is the one in the Quick Answer: DeepSeek V4-Pro 0424 is the pick on current neutral numbers, Kimi K2.6 and GLM-5.1 are tied behind it, and the gaps are small enough that licence, context and self-host constraints still decide most real decisions. The broader lesson from the rebase is worth keeping: a benchmark score with no version stamp is not evidence, and a trio that looked like the open-weights frontier in April 2026 is a tier below it by October.
If you are building or operating agentic coding infrastructure on these open-weights models and want senior engineers who have shipped it in production, Codersera matches you with vetted remote developers experienced with self-hosted LLM serving, agent harnesses, and cost-control tooling. We run a risk-free trial so you can validate technical fit before committing.
FAQ
Is Kimi K2.6 better than DeepSeek V4 for coding?
On current neutral data, DeepSeek V4-Pro. It leads Artificial Analysis's Intelligence Index v4.3.2 (read 5 October 2026) at 30.45 against Kimi K2.6's 26.98, costs $0.12 per index task against K2.6's $0.80, generates at 107.8 tokens/second against 59.6, and also wins AA's GDPval-AA agentic benchmark (1554 vs 1484). Vendor SWE-bench Verified numbers (80.2% vs 80.6%) remain effectively tied. Note that older write-ups — including an earlier version of this page — had Kimi ahead on the composite; that was the superseded v4.1.1-era index, and there is no conversion between index versions.
Which is the cheapest of the three?
It depends on whether you mean per token or per finished task. Per token, DeepSeek V4-Flash: $0.14 in / $0.28 out, roughly an order of magnitude below Kimi K2.6 ($0.95/$4.00) and GLM-5.1 ($1.40/$4.40), and about 3× below V4-Pro's $0.435/$0.87. Per completed task, Artificial Analysis measures them level — $0.11 for Flash against $0.12 for V4-Pro on Intelligence Index v4.3.2 — because Flash writes more tokens to get there. Kimi K2.6 costs $0.80 per task and GLM-5.1 $1.02. So V4-Pro is the value pick on the API and Flash is the value pick when you are paying for GPUs or capping output length yourself.
Do all three have truly open licenses?
DeepSeek V4 and GLM-5.1 are released under a clean, unmodified MIT license — commercial use, modification, and redistribution with no usage restrictions. Kimi K2.6 uses a Modified MIT license that adds a visible-attribution requirement for very large deployments. All three publish full weights on Hugging Face, so all three are genuinely self-hostable; the Kimi clause is the only license asterisk among them.
Which has the largest context window?
DeepSeek V4 — both V4-Pro and V4-Flash expose a 1M-token context window. Kimi K2.6 offers 256K and GLM-5.1 offers 200K. For whole-repository agentic tasks the DeepSeek 1M window is a real advantage, but remember that filling a 1M window costs input tokens every turn, so a smaller window you do not over-fill can be cheaper in practice.
Why do vendor benchmarks disagree with neutral ones?
Vendors run benchmarks with their own scaffolds, retry budgets, and effort/thinking settings, each tuned to show the model at its best. Independent harnesses like Artificial Analysis apply one fixed methodology to every model, which is why their numbers are comparable across models and vendor numbers are not. Trust vendor numbers to confirm a model is frontier-class; trust neutral numbers to rank models against each other.
Which is easiest to self-host?
DeepSeek V4-Flash, because at 284B total parameters it is by far the smallest of the four model variants while still shipping a 1M context and a clean MIT license. GLM-5.1's FP8 checkpoint needs roughly 8x H100 80GB; Kimi K2.6 wants 8x H200-class at full precision (though native INT4 quantization-aware training lets it run on 4x H100 at reduced context). All three support vLLM and SGLang.
Which one should a team standardise on in 2026?
If you need one default from this generation: DeepSeek V4-Pro 0424. It leads the trio on Intelligence Index v4.3.2 (30.45), is cheapest per completed task ($0.12), fastest (107.8 tok/s), MIT-licensed and 1M-context. Add V4-Flash alongside it only if you self-host or need a hard per-token cap, and pick GLM-5.1 instead if an unmodified MIT licence plus independent coding-Elo evidence is a compliance requirement. Before standardising, though, check the newer open-weights releases — MiMo-V2.6-Pro (46.32), GLM-5.3 (44.78) and Kimi K3 (43.59) all score well above everything in this table on the same index.
Does GLM 5.2 change this comparison?
Not immediately. GLM 5.2 shipped on June 13, 2026 with a 1M-token context window and MIT-licensed open weights (the week after launch), but Zhipu has not published benchmarks for 5.2 yet — so on independently verifiable numbers, GLM-5.1 remains the model in this table. If 5.2 holds 5.1's SWE-Bench Pro and Terminal-Bench 2.0 gains while adding the 1M window, it becomes the open-weights model with the most usable context. We'll update this comparison once independent third-party numbers land. Z.ai has since moved on again — GLM-5.3 landed on August 14, 2026, with no open weights and no public API at launch, so it carries no independent benchmark numbers either.