Kimi K3: Moonshot AI’s 2.8T Open-Weight Model — Release, Specs & Pricing (2026)

Kimi K3 is Moonshot AI's 2.8-trillion-parameter open-weight model. Pricing (API, OpenRouter, free kimi.com tier), specs, license, and Opus 5 comparison.

Quick answer. Kimi K3 is Moonshot AI's frontier model, released July 16, 2026. It's a 2.8-trillion-parameter open-weight, multimodal reasoning model — the largest open-weight model shipped to date — with a 1-million-token context window, an always-on "thinking mode," and API pricing of $3 per million input tokens and $15 per million output tokens ($0.30 on cache hits); it's free to use on kimi.com. The full open weights shipped July 27, 2026, under the modified-MIT "Kimi K3 License," and as of August 20, 2026, K3 is tied for the top open-weights score on the Artificial Analysis Intelligence Index (60).

Moonshot AI shipped Kimi K3 on July 16, 2026, and it immediately reset expectations for what an open-weight model can do. At 2.8 trillion parameters it is the largest open-weight model released so far, and on independent testing it ranks as the top open-weights model: on the Artificial Analysis Intelligence Index (v4.1.1, August 20, 2026) K3 scores 60, tied with GLM-5.3 and within three points of Claude Opus 5 (max, 63). For anyone building on open models, K3 is the first Chinese release that competes with the top U.S. systems on capability rather than just on price.

This guide covers what Kimi K3 actually is: the architecture, the specs, how to access it, what it costs, and where it fits. For the full head-to-head numbers against Fable 5, GPT-5.6 Sol, and Opus 5, see our companion Kimi K3 benchmarks comparison.

What is Kimi K3?

Kimi K3 is Moonshot AI's most capable model to date — a mixture-of-experts (MoE) large language model with 2.8 trillion total parameters, native multimodal (visual) understanding, and a one-million-token context window. Unlike the K2 series that preceded it, K3 ships with reasoning always on: Moonshot calls it "thinking mode," and the model reasons through problems by default rather than needing a separate reasoning variant.

It is a true open-weight release. The model went live via Kimi's apps and API on July 16, 2026, and the full weights landed on Hugging Face on July 27, 2026 — which is what makes the 2.8T parameter count so notable: nothing this large had been released with open weights before. The repo has logged over 2.3 million downloads in under a month (as of August 20, 2026), and Moonshot published the full K3 technical report the same day the weights dropped.

Kimi K3 key specifications

SpecificationKimi K3
Total parameters2.8 trillion
ArchitectureMixture-of-Experts (896 experts, 16 active per token)
Context window1,048,576 tokens (1M)
ModalityMultimodal (text + native visual understanding)
ReasoningAlways-on "thinking mode"
WeightsOpen weights (released July 27, 2026 — modified-MIT "Kimi K3 License")
Input price$3.00 / million tokens ($0.30 on cache hit)
Output price$15.00 / million tokens
ReleasedJuly 16, 2026

How does the Kimi K3 architecture work?

K3's headline efficiency trick is that it is enormous on paper but sparse in practice. Of its 896 experts, only 16 are activated for any given token — roughly 1.8% of the pool — so the compute cost of a forward pass is far lower than the 2.8T parameter count suggests. That is what lets a model this size run at a price comparable to mid-range Western models.

Two architectural innovations developed in-house at Moonshot do the heavy lifting:

  • Kimi Delta Attention — a hybrid linear attention mechanism that Moonshot reports delivers up to 6.3x faster decoding, which is a large part of how K3 stays usable across a full 1M-token context.
  • Attention Residuals — described as a drop-in replacement for standard residual connections. Moonshot reports a roughly 25% training-efficiency gain for under 2% additional compute overhead, which helped make training a 2.8T model economically feasible under China's compute constraints.

How much does Kimi K3 cost?

Kimi K3 pricing has three layers: the official Moonshot API, third-party hosts on OpenRouter, and the kimi.com consumer plans — where K3 is actually free to use. Here is the full picture as of August 20, 2026.

Official Kimi API pricing

The official API rate is $3.00 per million input tokens and $15.00 per million output tokens, with a $0.30 cache-hit input rate — the same tier as Anthropic's Claude Sonnet series, and unchanged since launch. Its predecessor, Kimi K2.6, runs roughly $0.95 in / $4 out, so K3 is a substantial step up.

Kimi K3 (official API)Rate per 1M tokens
Input (cache hit)$0.30
Input (cache miss)$3.00
Output$15.00
Context window1,048,576 tokens

Two details matter for real bills. First, there is no off-peak discount — unlike DeepSeek's peak/off-peak scheme, K3's rates are flat around the clock, and no batch tier is listed. Second, automatic context caching is standard, and Moonshot claims the official API exceeds a 90% cache-hit rate in coding workloads (a vendor claim, tied to its Mooncake serving stack) — meaning effective input cost in agent loops trends toward the $0.30 rate rather than $3.00.

OpenRouter and third-party pricing

Because the weights are open, a dozen-plus providers serve K3 through OpenRouter. Price competition is thin: as of August 20, 2026, the cheapest listing is Sail Research at $2.60 in / $13.00 out (fp4, slightly clipped context) — only about 13% below official — while most providers (Fireworks, Together, DigitalOcean, DeepInfra, and Moonshot's own first-party endpoint) sit at or near the official $3 / $15. K3's native weights are MXFP4, so heavily quantized cut-price copies don't really undercut the list rate.

Kimi.com plans: is Kimi K3 free?

Yes — K3 is free to use on kimi.com and the Kimi apps on the free Adagio tier, with limits (one concurrent agent task, two scheduled tasks). Paid plans run $19 to $199 per month and buy concurrency, priority, Kimi Code, agent swarms, and — from the Allegro tier up — 1M-token "extra long chat." One caveat: as of August 20, 2026, kimi.com carries a banner that new membership plans are coming soon, with Kimi and Kimi Code benefits to be separated (existing subscribers unaffected) — treat this table as a snapshot.

PlanMonthlyK3-relevant limits
Adagio (free)$0K3 access; 1 concurrent agent task; 2 scheduled tasks
Moderato$192 concurrent tasks, 4x priority queue, Kimi Code
Allegretto$39Goal mode, Kimi Claw, 4-subagent swarm
Allegro$79–994 concurrent tasks, 1M-token "extra long chat", 8-subagent swarm
Vivace$19910x agent credits, 25 scheduled tasks

Demand has been real enough that Moonshot briefly suspended new subscriptions on July 19, 2026 to protect capacity (since resumed).

What does K3 cost per task in practice?

Commentators read K3's launch pricing as the end of "super-cheap Chinese AI": Moonshot now prices on capability rather than undercutting. But on a per-task basis K3 still comes out cheaper than Western frontier models — Artificial Analysis measures about $0.84 per task for K3 versus $1.23 for GPT-5.6 Sol (max) and $2.34 for Claude Opus 5 (max) (AA v4.1.1, August 20, 2026) — because its answers tend to be token-efficient.

How to access Kimi K3

  1. Kimi apps: K3 is live on Kimi.com, the Kimi mobile apps, the Kimi Work desktop client, and Kimi Code.
  2. API: Available directly from Moonshot and through aggregators such as OpenRouter. The API's reasoning_effort parameter now accepts low, high, and max (default max) — the lower levels shipped after launch — alongside tool_choice constraints and dynamically loaded tools.
  3. Open weights: The full weights are live on Hugging Face (released July 27, 2026), so you can self-host or run K3 through open-model providers — subject to the hardware a 2.8T MoE needs (MXFP4-native, roughly 1.5TB to host).

If you're comparing open options more broadly, our Kimi K2.6 complete guide covers the prior generation, and the benchmarks piece below shows exactly where K3 sits against the closed frontier.

What's new since the Kimi K3 launch?

K3's first month brought a steady stream of updates (all dates 2026):

  • July 20 — Kimi Work: Moonshot launched its desktop work product (widgets + dashboards), another surface where K3 runs.
  • July 27 — open weights + technical report: the full 2.8T weights landed on Hugging Face (2.3M+ downloads by August 20) alongside the K3 technical report detailing Kimi Delta Attention, Attention Residuals, and the MXFP4 quantization-aware training recipe.
  • July 29 — K3-256k: a 256k-context K3 variant for Kimi Code — a leaner serving tier for coding workloads.
  • Reasoning controls: reasoning_effort gained low and high levels (only max existed at launch).
  • Leaderboard reshuffle: with Claude Opus 5, Grok 4.6, and GLM-5.3 all landing after K3, Artificial Analysis' rescaled v4.1.1 index now puts K3 at 60 — tied with GLM-5.3 as the top open-weights model (as of August 20, 2026).

Kimi K3 vs Claude Opus 5: which is better?

Anthropic shipped Claude Opus 5 on July 24, 2026 — eight days after K3 — and it is now K3's closest frontier comparison. On the Artificial Analysis Intelligence Index (v4.1.1, August 20, 2026), Opus 5 (max) scores 63 to K3 (max)'s 60, but K3 is 2.8x cheaper per measured task ($0.84 vs $2.34) and roughly matches the mid-effort Opus 5 (high) tier at 61. On list price, K3's $3 / $15 undercuts Opus 5's $5 / $25 by 40% on both sides.

On coding, the two are effectively tied: the independent Vals Index (updated August 12, 2026) has Claude Fable 5 at 75.14%, Opus 5 at 74.82%, and K3 at 74.70% — a 0.12-point gap. K3's weak spot is raw output speed (39 vs 59 tokens/second), and Moonshot's own report concedes a "noticeable gap in user experience" versus Claude Fable 5 and GPT-5.6 Sol. Pick Opus 5 for the highest-stakes agentic work and faster generation; pick K3 for open weights, native vision at a lower price, the 1M-token context, and per-task economics.

Who should use Kimi K3?

  • Engineering teams building on open weights — K3 is the first open-weight model that credibly competes with GPT-5.6 Sol and Claude on frontier coding tasks (it took #1 on the Frontend Code Arena).
  • Long-context workloads — repository-scale code navigation, large-document analysis, and log-heavy debugging benefit from the 1M-token window and strong long-context retention.
  • Cost-sensitive agentic pipelines — the token-efficiency means real per-task costs can undercut Western frontier models even at a similar per-token price.

One caveat worth knowing: Moonshot's own technical report concedes a remaining user-experience gap versus Claude Fable 5 and GPT-5.6 Sol, plus a tendency toward "excessive proactiveness" on ambiguous tasks — it recommends pinning explicit constraints in AGENTS.md-style instructions. As with any frontier model, keep verification in the loop for factual work.

Part of our AI models series. For the complete picture on Moonshot's lineup, read the Kimi complete guide, and see the full benchmark breakdown in Kimi K3 benchmarks vs Fable 5, Opus 5 & GPT-5.6 Sol.

Frequently asked questions

When was Kimi K3 released?

Moonshot AI released Kimi K3 on July 16, 2026, via its apps and API. The full open weights were published on Hugging Face on July 27, 2026.

How many parameters does Kimi K3 have?

Kimi K3 has 2.8 trillion total parameters, making it the largest open-weight model released to date. It uses a mixture-of-experts design with 896 experts, of which 16 are active per token.

Is Kimi K3 open source?

Kimi K3 is open-weight. The full weights shipped on July 27, 2026, under the "Kimi K3 License" — a modified MIT license allowing free commercial use, modification, distribution, and fine-tuning. The one condition beyond attribution: if you operate a "Model as a Service" business with over $20M in revenue across any 12 months, you must sign a separate agreement with Moonshot before commercial use. Self-hosting is otherwise unrestricted, subject to the substantial hardware a 2.8T MoE requires.

How much does Kimi K3 cost?

Kimi K3 is priced at $3 per million input tokens and $15 per million output tokens, with a $0.30 cache-hit input rate — roughly the Claude Sonnet tier and a notable increase over Kimi K2.6.

What is Kimi K3's context window?

Kimi K3 supports a 1-million-token context window, and it retains long-context performance well: on a 1M-token evaluation with no context management it scored 90.4 (a vendor-run figure).

Is Kimi K3 better than GPT-5.6 or Claude?

On the Artificial Analysis Intelligence Index (v4.1.1, as of August 20, 2026), K3 (max) scores 60 — behind Claude Opus 5 (max, 63), Claude Fable 5 (62), and GPT-5.6 Sol (max, 61), and tied with GLM-5.3 as the top open-weights model. It leads on the Frontend Code Arena and is markedly cheaper per measured task. See our full benchmark comparison for the details.

Is Kimi K3 better than Claude Opus 5?

Claude Opus 5 (max) leads Kimi K3 (max) 63 to 60 on the Artificial Analysis Intelligence Index (v4.1.1, August 20, 2026), but K3 is about 2.8x cheaper per measured task ($0.84 vs $2.34) and the two are near-tied on the Vals coding index (74.70% vs 74.82%). Opus 5 is the pick for the highest-stakes agentic work; K3 wins on open weights, list price ($3/$15 vs $5/$25), and 1M-token context.

The bottom line

Kimi K3 is the moment open weights caught up to the closed frontier. It won't top every leaderboard — Claude Opus 5, Fable 5, and GPT-5.6 Sol still edge it on the general-intelligence indexes — but a 2.8T open-weight model that ties for the top open-weights score, wins the frontend coding arena, near-ties Opus 5 on the Vals coding index, and holds a full 1M-token context is a genuine milestone. If your stack is built on open models, K3 is the first one you can reach for without a capability compromise.

Building something on frontier models and need engineers who move at this pace? Codersera connects you with vetted remote developers who ship with the latest AI tooling. Extend your engineering team without the hiring risk.