Kimi K3: Moonshot AI’s 2.8T Open-Weight Model — Release, Specs & Pricing (2026)

Kimi K3 is Moonshot AI's 2.8-trillion-parameter open-weight model. Pricing (API, OpenRouter, free kimi.com tier), specs, license, and Opus 5 comparison.

Quick answer. Kimi K3 is Moonshot AI's frontier model, released July 16, 2026. It's a 2.8-trillion-parameter open-weight, multimodal reasoning model — the largest open-weight model shipped to date — with a 1-million-token context window, an always-on "thinking mode," and API pricing of $3 per million input tokens and $15 per million output tokens ($0.30 on cache hits); it's free to use on kimi.com. The full open weights shipped July 27, 2026, under the modified-MIT "Kimi K3 License." On Artificial Analysis's rebased Intelligence Index v4.3.2 (read October 5, 2026) K3 scores 44 — third among open-weight models, behind MiMo-V2.6-Pro (46) and GLM-5.3 (45).

Moonshot AI shipped Kimi K3 on July 16, 2026, and it immediately reset expectations for what an open-weight model can do. At 2.8 trillion parameters it is the largest open-weight model released so far. It was the joint top-scoring open-weight model on independent testing through the summer — but that is no longer true, and the reason is worth understanding before you read any benchmark number on this page. Artificial Analysis rebased its Intelligence Index from v4.1.1 to v4.3.2, which is a reset rather than a drift: scores on the two versions are not comparable. Re-read on October 5, 2026 under v4.3.2, K3 (max) scores 44. That places it third among open weights — behind Xiaomi's MIT-licensed MiMo-V2.6-Pro (46) and Z.ai's GLM-5.3 (45) — and seven points behind Claude Opus 5 (max, 51), whose own successor Opus 5.5 now leads the whole index at 58. K3 remains the first Chinese release that competes with top U.S. systems on capability rather than price alone; it is no longer the open-weights leader.

This guide covers what Kimi K3 actually is: the architecture, the specs, how to access it, what it costs, and where it fits. For the full head-to-head numbers against Fable 5, GPT-5.6 Sol, and Opus 5, see our companion Kimi K3 benchmarks comparison.

What is Kimi K3?

Kimi K3 is Moonshot AI's most capable model to date — a mixture-of-experts (MoE) large language model with 2.8 trillion total parameters, native multimodal (visual) understanding, and a one-million-token context window. Unlike the K2 series that preceded it, K3 ships with reasoning always on: Moonshot calls it "thinking mode," and the model reasons through problems by default rather than needing a separate reasoning variant.

It is a true open-weight release. The model went live via Kimi's apps and API on July 16, 2026, and the full weights landed on Hugging Face on July 27, 2026 — which is what makes the 2.8T parameter count so notable: nothing this large had been released with open weights before. The repo has logged over 2.3 million downloads in under a month (as of August 20, 2026), and Moonshot published the full K3 technical report the same day the weights dropped.

Kimi K3 key specifications

SpecificationKimi K3
Total parameters2.8 trillion
ArchitectureMixture-of-Experts (896 experts, 16 active per token)
Context window1,048,576 tokens (1M)
ModalityMultimodal (text + native visual understanding)
ReasoningAlways-on "thinking mode"
WeightsOpen weights (released July 27, 2026 — modified-MIT "Kimi K3 License")
Input price$3.00 / million tokens ($0.30 on cache hit)
Output price$15.00 / million tokens
ReleasedJuly 16, 2026

How does the Kimi K3 architecture work?

K3's headline efficiency trick is that it is enormous on paper but sparse in practice. Of its 896 experts, only 16 are activated for any given token — roughly 1.8% of the pool — so the compute cost of a forward pass is far lower than the 2.8T parameter count suggests. That is what lets a model this size run at a price comparable to mid-range Western models.

Two architectural innovations developed in-house at Moonshot do the heavy lifting:

  • Kimi Delta Attention — a hybrid linear attention mechanism that Moonshot reports delivers up to 6.3x faster decoding, which is a large part of how K3 stays usable across a full 1M-token context.
  • Attention Residuals — described as a drop-in replacement for standard residual connections. Moonshot reports a roughly 25% training-efficiency gain for under 2% additional compute overhead, which helped make training a 2.8T model economically feasible under China's compute constraints.

How much does Kimi K3 cost?

Kimi K3 pricing has three layers: the official Moonshot API, third-party hosts on OpenRouter, and the kimi.com consumer plans — where K3 is actually free to use. Here is the full picture as of August 20, 2026.

Official Kimi API pricing

The official API rate is $3.00 per million input tokens and $15.00 per million output tokens, with a $0.30 cache-hit input rate — the same tier as Anthropic's Claude Sonnet series, and unchanged since launch. Its predecessor, Kimi K2.6, runs roughly $0.95 in / $4 out, so K3 is a substantial step up.

Kimi K3 (official API)Rate per 1M tokens
Input (cache hit)$0.30
Input (cache miss)$3.00
Output$15.00
Context window1,048,576 tokens

Two details matter for real bills. First, there is no off-peak discount — unlike DeepSeek's peak/off-peak scheme, K3's rates are flat around the clock, and no batch tier is listed. Second, automatic context caching is standard, and Moonshot claims the official API exceeds a 90% cache-hit rate in coding workloads (a vendor claim, tied to its Mooncake serving stack) — meaning effective input cost in agent loops trends toward the $0.30 rate rather than $3.00.

OpenRouter and third-party pricing

Because the weights are open, a dozen-plus providers serve K3 through OpenRouter. Price competition is thin: as of August 20, 2026, the cheapest listing is Sail Research at $2.60 in / $13.00 out (fp4, slightly clipped context) — only about 13% below official — while most providers (Fireworks, Together, DigitalOcean, DeepInfra, and Moonshot's own first-party endpoint) sit at or near the official $3 / $15. K3's native weights are MXFP4, so heavily quantized cut-price copies don't really undercut the list rate.

Kimi.com plans: is Kimi K3 free?

Yes — K3 is free to use on kimi.com and the Kimi apps on the free Adagio tier, with limits (one concurrent agent task, two scheduled tasks). Paid plans run $19 to $199 per month and buy concurrency, priority, Kimi Code, agent swarms, and — from the Allegro tier up — 1M-token "extra long chat." One caveat: as of August 20, 2026, kimi.com carries a banner that new membership plans are coming soon, with Kimi and Kimi Code benefits to be separated (existing subscribers unaffected) — treat this table as a snapshot.

PlanMonthlyK3-relevant limits
Adagio (free)$0K3 access; 1 concurrent agent task; 2 scheduled tasks
Moderato$192 concurrent tasks, 4x priority queue, Kimi Code
Allegretto$39Goal mode, Kimi Claw, 4-subagent swarm
Allegro$79–994 concurrent tasks, 1M-token "extra long chat", 8-subagent swarm
Vivace$19910x agent credits, 25 scheduled tasks

Demand has been real enough that Moonshot briefly suspended new subscriptions on July 19, 2026 to protect capacity (since resumed).

What does K3 cost per task in practice?

Commentators read K3's launch pricing as the end of "super-cheap Chinese AI": Moonshot now prices on capability rather than undercutting. On a per-task basis that reading has aged well, and the picture is more mixed than it looked in August. Under Artificial Analysis's Intelligence Index v4.3.2 (read October 5, 2026), the cost to complete one index task is $2.00 for K3 (max), $1.99 for GPT-5.6 Sol (max) and $5.86 for Claude Opus 5 (max). So K3 is roughly a third of the price of Opus 5 per task, but it is now level with GPT-5.6 Sol rather than cheaper than it — and GPT-5.6 Sol scores higher on the index (47 against 44). The token-efficiency advantage has also gone: K3 burned 160M output tokens completing the v4.3.2 index, against 140M for Opus 5 and 90M for GPT-5.6 Sol.

How to access Kimi K3

  1. Kimi apps: K3 is live on Kimi.com, the Kimi mobile apps, the Kimi Work desktop client, and Kimi Code.
  2. API: Available directly from Moonshot and through aggregators such as OpenRouter. The API's reasoning_effort parameter now accepts low, high, and max (default max) — the lower levels shipped after launch — alongside tool_choice constraints and dynamically loaded tools.
  3. Open weights: The full weights are live on Hugging Face (released July 27, 2026), so you can self-host or run K3 through open-model providers — subject to the hardware a 2.8T MoE needs (MXFP4-native, roughly 1.5TB to host).

If you're comparing open options more broadly, our Kimi K2.6 complete guide covers the prior generation, and the benchmarks piece below shows exactly where K3 sits against the closed frontier.

What's new since the Kimi K3 launch?

K3's first month brought a steady stream of updates (all dates 2026):

  • July 20 — Kimi Work: Moonshot launched its desktop work product (widgets + dashboards), another surface where K3 runs.
  • July 27 — open weights + technical report: the full 2.8T weights landed on Hugging Face (2.3M+ downloads by August 20) alongside the K3 technical report detailing Kimi Delta Attention, Attention Residuals, and the MXFP4 quantization-aware training recipe.
  • July 29 — K3-256k: a 256k-context K3 variant for Kimi Code — a leaner serving tier for coding workloads.
  • Reasoning controls: reasoning_effort gained low and high levels (only max existed at launch).
  • Leaderboard reshuffle: Artificial Analysis has since rebased its Intelligence Index to v4.3.2, and on that scale K3 (max) scores 44 — third among open weights, behind MiMo-V2.6-Pro (46) and GLM-5.3 (45), with Claude Opus 5.5 leading the overall index at 58 (read October 5, 2026). The v4.1.1 figure of 60 that circulated in August is on a different scale and is not convertible to the new one.

Kimi K3 vs Claude Opus 5: which is better?

Anthropic shipped Claude Opus 5 on July 24, 2026 — eight days after K3 — and it was K3's closest frontier comparison for most of the summer. On Artificial Analysis's Intelligence Index v4.3.2 (read October 5, 2026), Opus 5 (max) scores 51 to K3 (max)'s 44 — a seven-point gap, wider than the three points the older v4.1.1 numbers showed, because the rebased index spreads the top of the field out further. K3's cost advantage survives the rebase intact: $2.00 against $5.86 per index task, so about 2.9x cheaper. On list price, K3's $3 / $15 undercuts Opus 5's $5 / $25 by 40% on both sides. One caveat on framing: Artificial Analysis now marks Opus 5 deprecated in favour of Claude Opus 5.5, which leads the index at 58 — so if you are choosing today, Opus 5.5 rather than Opus 5 is the real Anthropic comparison.

On coding, the two are effectively tied: the independent Vals Index (updated August 12, 2026) has Claude Fable 5 at 75.14%, Opus 5 at 74.82%, and K3 at 74.70% — a 0.12-point gap. K3's weak spot is raw output speed (45.0 against 56.0 tokens per second on Artificial Analysis's October 2026 measurements), and Moonshot's own report concedes a "noticeable gap in user experience" versus Claude Fable 5 and GPT-5.6 Sol. Pick Opus 5 for the highest-stakes agentic work and faster generation; pick K3 for open weights, native vision at a lower price, the 1M-token context, and per-task economics.

Who should use Kimi K3?

  • Engineering teams building on open weights — K3 is the first open-weight model that credibly competes with GPT-5.6 Sol and Claude on frontier coding tasks (it took #1 on the Frontend Code Arena).
  • Long-context workloads — repository-scale code navigation, large-document analysis, and log-heavy debugging benefit from the 1M-token window and strong long-context retention.
  • Cost-sensitive agentic pipelines — the token-efficiency means real per-task costs can undercut Western frontier models even at a similar per-token price.

One caveat worth knowing: Moonshot's own technical report concedes a remaining user-experience gap versus Claude Fable 5 and GPT-5.6 Sol, plus a tendency toward "excessive proactiveness" on ambiguous tasks — it recommends pinning explicit constraints in AGENTS.md-style instructions. As with any frontier model, keep verification in the loop for factual work.

Part of our AI models series. For the complete picture on Moonshot's lineup, read the Kimi complete guide, and see the full benchmark breakdown in Kimi K3 benchmarks vs Fable 5, Opus 5 & GPT-5.6 Sol.

Frequently asked questions

When was Kimi K3 released?

Moonshot AI released Kimi K3 on July 16, 2026, via its apps and API. The full open weights were published on Hugging Face on July 27, 2026.

How many parameters does Kimi K3 have?

Kimi K3 has 2.8 trillion total parameters, making it the largest open-weight model released to date. It uses a mixture-of-experts design with 896 experts, of which 16 are active per token.

Is Kimi K3 open source?

Kimi K3 is open-weight. The full weights shipped on July 27, 2026, under the "Kimi K3 License" — a modified MIT license allowing free commercial use, modification, distribution, and fine-tuning. The one condition beyond attribution: if you operate a "Model as a Service" business with over $20M in revenue across any 12 months, you must sign a separate agreement with Moonshot before commercial use. Self-hosting is otherwise unrestricted, subject to the substantial hardware a 2.8T MoE requires.

How much does Kimi K3 cost?

Kimi K3 is priced at $3 per million input tokens and $15 per million output tokens, with a $0.30 cache-hit input rate — roughly the Claude Sonnet tier and a notable increase over Kimi K2.6.

What is Kimi K3's context window?

Kimi K3 supports a 1-million-token context window, and it retains long-context performance well: on a 1M-token evaluation with no context management it scored 90.4 (a vendor-run figure).

Is Kimi K3 better than GPT-5.6 or Claude?

No, not on the general-intelligence index. On Artificial Analysis's Intelligence Index v4.3.2 (read October 5, 2026), K3 (max) scores 44 — behind Claude Opus 5.5 (max, 58, the current index leader), Claude Opus 5 (max, 51), Claude Fable 5 (max, 50) and GPT-5.6 Sol (max, 47). Among open-weight models it is third, behind MiMo-V2.6-Pro (46) and GLM-5.3 (45). It leads on the Frontend Code Arena, and it is far cheaper per index task than the Anthropic models ($2.00 against $5.86 for Opus 5), though now level with GPT-5.6 Sol at $1.99. See our full benchmark comparison for the details.

Is Kimi K3 better than Claude Opus 5?

Claude Opus 5 (max) leads Kimi K3 (max) 51 to 44 on Artificial Analysis's Intelligence Index v4.3.2 (read October 5, 2026), but K3 is about 2.9x cheaper per index task ($2.00 against $5.86) and the two are near-tied on the Vals coding index (74.70% vs 74.82%). Opus 5 is the pick for the highest-stakes agentic work — though Anthropic has since superseded it with Opus 5.5, which scores 58. K3 wins on open weights, list price ($3/$15 vs $5/$25), and 1M-token context.

The bottom line

Kimi K3 was the moment open weights caught up to the closed frontier, and it is worth being precise about what has changed since. It does not top the open-weights leaderboard any more: on Artificial Analysis's Intelligence Index v4.3.2 it sits at 44, behind MiMo-V2.6-Pro (46) and GLM-5.3 (45), with Claude Opus 5.5 leading everything at 58. What has not changed is the rest of the case — a 2.8T open-weight model that wins the frontend coding arena, near-ties Opus 5 on the Vals coding index, holds a full 1M-token context, and costs about a third of Opus 5 per measured task. Reach for K3 when you want a large open-weight model with a genuine 1M window and Anthropic-adjacent coding quality at a third of the price. If what you want is simply the highest open-weights score, MiMo-V2.6-Pro and GLM-5.3 are both ahead of it now, and both are worth benchmarking against your own workload before you commit.

Building something on frontier models and need engineers who move at this pace? Codersera connects you with vetted remote developers who ship with the latest AI tooling. Extend your engineering team without the hiring risk.