Quick answer. Kimi K3 scores 60 on the Artificial Analysis Intelligence Index v4.1.1 — the top open-weight result, but behind Claude Opus 5 (63), Claude Fable 5 (62) and GPT-5.6 Sol (61). It wins outright on frontend coding (#1 on LMArena's Frontend Code Arena, 1,679 Elo) and costs roughly 40% less per blended million tokens than Opus 5.
Moonshot AI released Kimi K3 on July 16, 2026 and shipped the full 2.8-trillion-parameter weights to Hugging Face on July 27. Since then the model has been scored by Moonshot itself, by Artificial Analysis, by Vals AI, by LMArena, and by at least one engineering team running its own private SWE-bench. Those five sources do not agree with each other — and the disagreement is the most useful thing on this page.
This is the benchmark-data reference for Kimi K3: every score we could verify against a primary source, what conditions it was measured under, and where the numbers stop being comparable. If you specifically want the K3-versus-Anthropic head-to-head, we have a dedicated deep dive: Kimi K3 vs Claude Opus 5.
What do Kimi K3's benchmark scores actually say?
Here is every verified figure in one place. Rows are benchmarks, columns are models.
| Benchmark | Kimi K3 | Claude Opus 5 | Claude Fable 5 | GPT-5.6 Sol | Claude Opus 4.8 |
|---|---|---|---|---|---|
| AA Intelligence Index v4.1.1 (max tier) | 60 | 63 | 62 | 61 | — |
| Vals Index v2 (Aug 19, 2026) | 57.8% | 67.2% | 66.0% | 63.7% | 60.9% |
| LMArena Frontend Code Arena (Elo) | 1,679 (#1) | not listed | 1,631 | 1,618 | 1,562 |
| SWE-bench Verified | not published | 96.0% | 95.5% | ~90%* | 88.6% |
| Terminal-Bench 2.1 | 88.3 | — | 88.0 | 88.8 | 84.6 |
| DeepSWE | 67.5 | — | 70.0 | 73.0 | 59.0 |
| FrontierSWE | 81.2 | — | 86.6 | 71.3 | 66.7 |
| SWE-Marathon | 42.0 | — | 35.0 | 39.0 | 40.0 |
| ProgramBench | 77.8 | — | 76.8 | 77.6 | 71.9 |
| Kimi Code Bench 2.0 | 72.9 | — | 76.9 | 64.8 | 71.7 |
| GPQA Diamond | 93.5 | — | 92.6 | 94.1 | 91.0 |
| HLE-Full (no tools) | 43.5 | — | 53.3 | 44.5 | 49.8 |
| GDPval-AA v2 (Elo) | 1,686 | — | 1,747 | 1,736 | 1,593 |
| AA-Briefcase (Elo) | 1,548 | — | 1,583 | 1,495 | 1,354 |
| BrowseComp | 91.2 | — | 88.0 | 90.4 | 84.3 |
| OSWorld 2.0 | 58.3 | — | 66.1 | 62.6 | 55.7 |
| List price (input / output per 1M) | $3 / $15 | $5 / $25 | $10 / $50 | $4 / $20 | $5 / $25 |
| Output speed (AA measured) | 37 tok/s | 56 tok/s | — | 72 tok/s | — |
How to read this table. The dashes are not missing data — they are the story. Everything from Terminal-Bench down to OSWorld comes from Moonshot's own Kimi K3 model card, which was published on July 16. Claude Opus 5 launched on July 24. Moonshot has never benchmarked K3 against Opus 5, so on the benchmarks Moonshot chose there is simply no head-to-head number. The only cross-vendor rows are the ones run by third parties: Artificial Analysis, Vals AI and LMArena.
Three more caveats that matter if you plan to cite any of this:
- Harnesses differ by design. Moonshot ran K3's coding benchmarks through its own Kimi Code harness while running the comparison models through Claude Code, Codex, or each vendor's official harness. An agentic score is a score for a model plus its scaffold, so a two-point gap on Terminal-Bench is well inside harness noise.
- All K3 figures are max-effort. Moonshot states that every result was obtained with reasoning effort set to
maxand temperature 1.0 (top-p 0.95 for single-step tasks, 1.0 for agentic ones). If you call K3 at a lower effort, you are not getting these numbers. - Anthropic and OpenAI report differently. SWE-bench Verified is the benchmark Anthropic and OpenAI lead with; Moonshot doesn't publish it for K3 at all. The GPT-5.6 Sol SWE-bench figure marked * is a third-party estimate, not an OpenAI-published number — treat it as approximate.
Where does Kimi K3 rank on the Artificial Analysis Intelligence Index?
At launch, Artificial Analysis scored K3 at 57 and placed it third overall, comparable to Opus 4.8 and GPT-5.5. That was July 17. Since then Opus 5, Grok 4.6 and GLM-5.3 all shipped and AA rescaled the index to v4.1.1. On the current version, K3 (max) scores 60 against Opus 5 (max) at 63, Fable 5 at 62 and GPT-5.6 Sol (max) at 61.
Here is a subtlety worth knowing before you quote a rank. AA scores models per reasoning-effort tier, and the ordering flips depending on which tier you look at. Compare max tiers and GPT-5.6 Sol edges K3, 61 to 60. Look at the aggregate leaderboard snapshot for August 22, which uses each model's default tier, and the order reverses: Kimi K3 at 59.7% sits above GPT-5.6 Sol at 58.9%, with Grok 4.6 (60.9%) slotting between K3 and Fable 5. Both statements are true. Neither is a clean "K3 beats Sol" claim.
What is not ambiguous: K3 is at or near the top of the open-weight field on this index, with GLM-5.3 (59.5%) as its only real peer at that tier. For the wider field, see our open-source LLM landscape guide.
How does Kimi K3 compare to GPT-5.6 Sol?
This is the closest matchup on the board, and it splits cleanly by task type.
GPT-5.6 Sol wins the reasoning and long-horizon coding rows. On Moonshot's own card, Sol takes GPQA Diamond (94.1 vs 93.5), DeepSWE (73.0 vs 67.5), Terminal-Bench 2.1 (88.8 vs 88.3), CritPt (32.3 vs 23.4) and OSWorld 2.0 (62.6 vs 58.3). It also leads on the Vals Index v2 by nearly six points, 63.7% to 57.8%. Sol generates roughly twice as fast — 72 tokens/second against K3's 37 in Artificial Analysis's measurements.
Kimi K3 wins the frontend, browsing and marathon rows. K3 takes FrontierSWE by ten points (81.2 vs 71.3), SWE-Marathon (42.0 vs 39.0), ProgramBench (77.8 vs 77.6), BrowseComp (91.2 vs 90.4), MCPMark-Verified (94.5 vs 92.9), AA-Briefcase Elo (1,548 vs 1,495) and Kimi Code Bench 2.0 (72.9 vs 64.8 — though that one is Moonshot's own benchmark, so weight it accordingly). On LMArena's Frontend Code Arena, K3 sits at 1,679 Elo against Sol's 1,618.
Pricing has moved recently and is worth re-checking: OpenAI now lists GPT-5.6 Sol at $4 input / $20 output per million tokens, described as promotional pricing available at least through November 21, 2026 — down from $5 / $30. That narrows K3's list-price advantage from roughly 2x to about 25%. Full tier breakdown in our GPT-5.6 Sol, Terra and Luna guide.
How does Kimi K3 compare to Claude Fable 5?
Fable 5 is Anthropic's top-tier model and it beats K3 nearly everywhere that isn't frontend code. On Moonshot's own card — a card written to make K3 look good — Fable 5 still wins HLE-Full (53.3 vs 43.5), FrontierSWE (86.6 vs 81.2), OSWorld 2.0 (66.1 vs 58.3), Kimi Code Bench 2.0 (76.9 vs 72.9), SciCode (60.2 vs 58.7), JobBench (57.4 vs 54.3), GDPval-AA v2 Elo (1,747 vs 1,686) and AA-Briefcase Elo (1,583 vs 1,548).
K3's wins over Fable 5 are narrower and more specific: SWE-Marathon (42.0 vs 35.0), Terminal-Bench 2.1 (88.3 vs 88.0), ProgramBench (77.8 vs 76.8), BrowseComp (91.2 vs 88.0), MCPMark-Verified (94.5 vs 87.4), GPQA Diamond (93.5 vs 92.6), Harvey Lab-AA (94.6 vs 93.6) and OmniDocBench (91.1 vs 89.8).
The headline K3-over-Fable-5 result is LMArena's Frontend Code Arena, where K3 took first place at 1,679 Elo against Fable 5's 1,631 and won six of the seven frontend domains — losing only Gaming, where Fable 5 keeps the edge. That is a genuine, third-party, human-preference win, and it is a 17-place jump from Kimi K2.6's #18.
The economics are not close. Anthropic lists Fable 5 at $10 input / $50 output per million tokens against K3's $3 / $15 — a 3.3x gap in both directions. If your workload is frontend generation, K3 wins on quality and costs a third as much. If it's hard multi-step agentic work, you are paying Fable 5's premium for a real capability gap.
How does Kimi K3 compare to Claude Opus 5?
Short version: Opus 5 leads on every cross-vendor index we can verify, and the gap is wider than the launch-week coverage suggested. On AA v4.1.1 it's 63 to 60. On the Vals Index v2 — a GDP-weighted blend of finance, coding and legal agentic tasks measured on August 19, 2026 — Opus 5 leads first place at 67.2% while K3 sits tenth at 57.8%, behind Opus 4.8, Sonnet 5, GPT-5.6 Terra, Gemini 3.7 Flash and Grok 4.6. Opus 5 also posts 96.0% on SWE-bench Verified and 79.2% on SWE-bench Pro, benchmarks K3 doesn't report.
K3's counter-arguments are open weights, price ($3/$15 vs $5/$25 list; $2.31 vs $3.85 per blended million on Artificial Analysis's 7:2:1 cache/input/output mix), a marginally larger context window (1,049k vs 1,000k), native vision, and a far lower time-to-first-token (3.09s vs 32.30s). Opus 5 answers with 56 tokens/second against K3's 37 once generation starts.
We've broken this matchup down properly — benchmark by benchmark, plus which one to actually put in your agent loop — in Kimi K3 vs Claude Opus 5. Opus 5's own scores are unpacked in our Opus 5 benchmarks explainer.
What does Kimi K3 actually cost per result?
Per-token price is the wrong unit for a reasoning model. What matters is cost to finish a task, and that depends on how many tokens the model burns getting there.
| Model | List input / output per 1M | Cache read | AA blended per 1M* |
|---|---|---|---|
| Kimi K3 | $3.00 / $15.00 | $0.30 | $2.31 |
| GPT-5.6 Sol | $4.00 / $20.00 | $0.40 | $4.35 |
| Claude Opus 5 | $5.00 / $25.00 | $0.50 | $3.85 |
| Claude Opus 4.8 | $5.00 / $25.00 | $0.50 | — |
| Claude Fable 5 | $10.00 / $50.00 | $1.00 | — |
*Artificial Analysis's blended figure assumes a 7:2:1 cache-hit / input / output token mix. GPT-5.6 Sol's list price is promotional at least through November 21, 2026.
Two independent cost-per-result measurements point the same way:
- Artificial Analysis measured K3's cost to run its full Intelligence Index at $0.94 per task at launch, against $1.04 for GPT-5.6 Sol and $1.80 for Opus 4.8. K3 consumed 132M output tokens across the index — 21% fewer than Kimi K2.6's 166M, so the generation-heavy reputation is improving.
- Superconductor ran a private SWE-bench built from its own merged Rails pull requests and found K3 matched Opus 4.8 on quality (~80%) at roughly 25% of the cost per ticket — the best open-weight result in their test.
There is a real counterweight, though. K3 took about 44 minutes per ticket in that same test — the slowest model they benchmarked, roughly twice Opus 4.8 and more than four times the GPT-5.6 family, which finished under ten minutes. If an engineer is waiting on the output, a 4x wall-clock penalty erases a 4x token saving. K3's economics work for asynchronous batch work: overnight refactors, bulk code review, migration passes. They work much less well for interactive pairing.
One more note on the open-weights argument. The weights are genuinely downloadable, but at 2.8T parameters the release is a 1.56 TB download, and the licence is Moonshot's own Kimi K3 License rather than MIT or Apache. For almost every team, "open weights" here means provider choice and price competition — thirteen providers currently serve K3, some below list at about $2.60 / $13.00 — not actually running it in your own rack.
Where does Kimi K3 lose?
A benchmark page that only lists wins isn't worth citing. Here is the honest ledger.
- Broad economic-task performance. Vals Index v2 puts K3 tenth at 57.8%, nearly ten points behind Opus 5 and below Opus 4.8, Sonnet 5, both smaller GPT-5.6 tiers, Gemini 3.7 Flash and Grok 4.6. Note that Vals restated its scoring in v2 — older figures around 74% that still circulate are from v1 and are not comparable to current numbers.
- Hard agentic work. GDPval-AA v2 has K3 at 1,686 Elo behind both Fable 5 (1,747) and GPT-5.6 Sol (1,736). AA's own run of the same benchmark scored K3 slightly lower at 1,668, which tells you how much Elo drift to expect between runs.
- Hardest reasoning. HLE-Full without tools: 43.5 against Fable 5's 53.3 and Opus 4.8's 49.8. CritPt: 23.4 against Sol's 32.3.
- Speed. 37 tokens/second measured by AA, against 56 for Opus 5 and 72 for GPT-5.6 Sol. Throughput also varies enormously by host: OpenRouter shows 22 to 133 tokens/second across providers, with 5 to 30 second latency.
- Factual reliability. On AA-Omniscience, K3's accuracy improved to 46% (from K2.6's 33%) but its hallucination rate rose to 51%. That is a model you keep verification around, not one you point at a knowledge base unsupervised.
- Reasoning-token overhead. Independent testers have hit requests where K3 spent 4,093 of 4,096 available completion tokens on reasoning and returned nothing usable. Set generous output budgets.
- No SWE-bench Verified. The industry's most-quoted agentic coding number is absent from Moonshot's card. K3 does appear on the official DeepSWE leaderboard at 67.3 under the mini-SWE-agent harness, close to its self-reported 67.5 — but that's a different benchmark from the one Anthropic and OpenAI lead with.
- Unpublished evaluation detail. Full prompts, evaluation code and end-to-end cost accounting are not public, so no one outside Moonshot can independently reproduce the card.
Which model should you pick for which job?
- Frontend and UI generation: Kimi K3. It's #1 on the Frontend Code Arena on human preference, and it costs a third of Fable 5.
- Highest-stakes agentic and long-horizon work: Claude Opus 5, with Fable 5 above it if budget is genuinely no object. The Vals and GDPval gaps are too wide to argue with.
- Interactive pairing where latency matters: GPT-5.6 Sol or Opus 5. K3's 37 tok/s and 44-minute ticket times are disqualifying for anything you sit and watch.
- Bulk asynchronous coding — overnight refactors, batch review, migrations: Kimi K3. Opus-4.8-class quality at roughly a quarter of the cost, and nobody is waiting.
- You need weights you control: Kimi K3 or GLM-5.3. K3 is stronger; note the licence is Moonshot's own, not a standard permissive one.
- Knowledge-heavy factual work: None of these unsupervised, and K3 least of all at a 51% hallucination rate.
Part of our AI models series. Read the Kimi K3 complete guide for architecture, hardware and access, the Kimi K3 vs Claude Opus 5 head-to-head, and DeepSeek V4 vs Kimi K3 for the open-weights matchup.
FAQ
Is Kimi K3 better than Claude Opus 5?
No, not on the cross-vendor indices. Opus 5 leads 63 to 60 on the Artificial Analysis Intelligence Index v4.1.1 and 67.2% to 57.8% on the Vals Index v2, and posts 96.0% on SWE-bench Verified, which K3 doesn't report. K3's advantages are open weights, a 40% lower blended token price, and a much faster time to first token.
Is Kimi K3 better than GPT-5.6 Sol?
It depends on the task. Sol leads GPQA Diamond, DeepSWE, Terminal-Bench 2.1 and the Vals Index, and generates roughly twice as fast. K3 leads FrontierSWE by ten points, plus SWE-Marathon, BrowseComp, MCPMark and the Frontend Code Arena. On the Artificial Analysis index the ranking flips depending on which reasoning-effort tier you compare.
Is Kimi K3 better than Claude Fable 5?
Only on frontend coding and a handful of tool-use benchmarks. K3 took first place on LMArena's Frontend Code Arena at 1,679 Elo against Fable 5's 1,631, winning six of seven domains. Fable 5 leads HLE-Full, FrontierSWE, OSWorld 2.0, GDPval-AA and AA-Briefcase — but costs $10/$50 per million tokens against K3's $3/$15.
Is Kimi K3 better than Claude Opus 4.8?
On most published benchmarks, yes. Moonshot's card shows K3 ahead of Opus 4.8 on Terminal-Bench 2.1 (88.3 vs 84.6), DeepSWE (67.5 vs 59.0), FrontierSWE (81.2 vs 66.7), GPQA Diamond (93.5 vs 91.0) and GDPval-AA Elo (1,686 vs 1,593). Independent testing found matching quality at roughly a quarter of the cost per ticket — but at about twice the wall-clock time.
Does Kimi K3 have a SWE-bench Verified score?
Moonshot does not publish one. Its card reports DeepSWE (67.5), FrontierSWE (81.2), SWE-Marathon (42.0), ProgramBench (77.8) and Terminal-Bench 2.1 (88.3) instead. K3 appears on the official DeepSWE leaderboard at 67.3 under the mini-SWE-agent harness. Anthropic and OpenAI lead with SWE-bench Verified, so no direct comparison on that benchmark exists.
What is the best open-weight model in 2026?
Kimi K3, by most measures — it leads the open-weight field on the Artificial Analysis Intelligence Index at 60, with GLM-5.3 (59.5%) its only close peer, and it's the strongest open-weight result in independent private SWE-bench testing. At 2.8T parameters and a 1.56 TB download it is also the least practical to self-host.
How much cheaper is Kimi K3 than the closed frontier?
List price is $3 input / $15 output per million tokens, against $4/$20 for GPT-5.6 Sol (promotional), $5/$25 for Claude Opus 5 and $10/$50 for Claude Fable 5. On Artificial Analysis's blended mix that works out to $2.31 per million for K3 versus $3.85 for Opus 5. Independent per-ticket testing found roughly a 4x saving against Opus 4.8.
Why do Kimi K3's rankings differ between leaderboards?
Three reasons. Different reasoning-effort tiers are scored separately, and K3 beats GPT-5.6 Sol on default tiers while losing on max. Different harnesses are used per model, so agentic scores measure model plus scaffold. And index versions get rescaled — Vals restated scores in v2 and Artificial Analysis rescaled to v4.1.1, so figures from July aren't comparable to August.
Can you actually run Kimi K3 yourself?
In principle. The full weights went up on Hugging Face on July 27, 2026 under Moonshot's own Kimi K3 License — not MIT or Apache. But the model is 2.8T parameters and a roughly 1.56 TB download, so for most teams the practical benefit of open weights is provider competition rather than local inference. Thirteen providers currently serve it.
The bottom line
Kimi K3 is the best open-weight model available in August 2026 and it is genuinely first on one meaningful third-party leaderboard — human-judged frontend code. It is not a frontier model. On the two cross-vendor indices that weight broad economic and agentic work, Artificial Analysis and Vals, it sits behind Claude Opus 5, Claude Fable 5 and GPT-5.6 Sol, and on Vals the gap to Opus 5 is close to ten points.
The decision rule that falls out of the data is about latency, not capability. K3 delivers Opus-4.8-class output at roughly a quarter of the cost, and takes roughly twice as long to get there. If a person is waiting, that trade is bad. If a queue is waiting, it's excellent — and it's the strongest version of that trade anyone has shipped with downloadable weights.
Building on models that change this fast? Extend your engineering team with vetted remote developers through Codersera.