Kimi K3 Pricing: API Costs, Plans & Real Bills (2026)

Kimi K3 costs $3 per 1M input tokens, $0.30 on a cache hit, and $15 per 1M output — flat across the full 1M-token context. Here are the verified rates, the rate-limit tiers, three worked cost examples, and how it prices against Claude Opus 5, GPT-5.6 and DeepSeek.

Quick answer. Kimi K3 costs $3.00 per 1M input tokens and $15.00 per 1M output tokens, dropping to $0.30 per 1M on a cache hit. One flat rate covers the full 1,048,576-token context — no long-context surcharge, no peak-hour multiplier, no separate thinking-token charge. Consumer plans run $19–$199 a month.

What does Kimi K3 cost per million tokens?

Moonshot AI publishes a single rate card for Kimi K3. There is no "mini", no "turbo", and no separate thinking-mode SKU — K3's reasoning is always on and its thinking tokens bill as ordinary output tokens. Reasoning depth is a request parameter (low, high, max), not a price tier.

Access routeInput /1MCached input /1MOutput /1MContextNotes
Kimi K3 — official API (kimi-k3)$3.00$0.30$15.001,048,576Flat across the whole window; prices exclude tax
OpenRouter — cheapest listing$2.60$0.29$13.00~1MSail Research; fp4, slightly clipped context
OpenRouter — typical provider$3.00$0.29$15.001,048,576Fireworks, Together, DeepInfra, Baseten, Moonshot's own endpoint
OpenRouter — most expensive$6.00—$22.50~1MMorph Fast; latency-optimised
Batch APINot available for K3 — Moonshot's 60%-of-list batch tier covers only K2.5, K2.6 and K2.7-code
$web_search built-in tool$0.005 per successful call, plus the returned results billed as normal input tokens

Rates in USD per 1M tokens, verified against Moonshot's K3 pricing page and OpenRouter on 23 August 2026.

Two things are unusual here. First, the cache-hit rate is a straight 90% discount with no cache-write premium and no storage fee — unlike Anthropic, which charges 1.25× to write a cache entry. Second, Moonshot does not tier by context length. Many Chinese labs do, and DeepSeek now runs a 2× peak-hour multiplier; K3 does neither. The $3 you pay on a 2,000-token prompt is the $3 you pay on a 900,000-token one.

What do you have to pay before the API will answer?

There are no free API credits. You top up a balance, and your recharge total determines your rate limits. The minimum is $1, and cumulative recharges reaching $5 earn a $5 voucher (vouchers themselves do not count toward the tier threshold).

TierCumulative rechargeConcurrencyRPMTPMTokens/day
Tier 0$113500,0001,500,000
Tier 1$10502002,000,000Unlimited
Tier 2$201005003,000,000Unlimited
Tier 3$1002005,0003,000,000Unlimited
Tier 4$1,0004005,0004,000,000Unlimited
Tier 5$3,0001,00010,0005,000,000Unlimited

The step that matters is $1 → $10. Tier 0 gives you 3 requests per minute and a 1.5M-token daily ceiling, which is a smoke-test allowance, not a workload. Ten dollars buys 50-way concurrency and removes the daily cap entirely. If you are evaluating K3 for an agent loop, put in $10 before you conclude it is slow.

What does Kimi K3 actually cost on real workloads?

List price tells you very little until you multiply it by a token profile. Three worked examples, arithmetic shown.

A coding-agent session over a repository

Assume 40 turns, each resending a growing context averaging 60,000 prompt tokens, with 1,500 output tokens per turn. That is 2.4M input and 60k output.

  • No caching: 2.4 × $3.00 = $7.20, plus 0.06 × $15.00 = $0.90 → $8.10
  • At a 90% cache-hit rate (Moonshot's stated figure for coding traffic on its Mooncake serving stack): 2.16 × $0.30 = $0.65, plus 0.24 × $3.00 = $0.72, plus $0.90 output → $2.27

Caching is worth 3.6× here because the prefix is enormous and stable. This is the workload K3's pricing is shaped for.

A document-analysis batch job

1,000 documents, ~50,000 tokens each, a shared 2,000-token system prompt, 800 tokens of output apiece.

  • Cached prefix: 999 × 2,000 = 2.0M at $0.30 → $0.60
  • Uncached document bodies: 50.0M at $3.00 → $150.00
  • Output: 0.8M at $15.00 → $12.00
  • Total ≈ $162.60, or $0.163 per document

Notice how little the cache saved: 0.4%. Caching rewards repeated prefixes, and a fan-out over distinct documents has almost none. Worse, K3 is excluded from Moonshot's batch tier, so there is no 40% discount to fall back on — a job you could halve on OpenAI or Anthropic's batch endpoints, you pay full freight for here.

A chat product at 10,000 messages a day

Assume 3,200 prompt tokens per message (system prompt plus a few turns of history) and 350 output tokens, with a realistic 60% cache-hit rate.

  • Cached input: 19.2M at $0.30 → $5.76
  • Uncached input: 12.8M at $3.00 → $38.40
  • Output: 3.5M at $15.00 → $52.50
  • $96.66/day ≈ $2,900/month, or $0.0097 per message

Output is 54% of that bill on 10% of the tokens. For chat, the lever is response length and reasoning effort — not context trimming.

Is Kimi K3 cheaper than Claude Opus 5, GPT-5.6 and Gemini?

On list price K3 sits mid-pack among frontier models and far above the discount Chinese tier. Cheaper-is-better is the wrong reading, so the last two columns give a capability-adjusted view using Artificial Analysis's Intelligence Index v4.3.2 (read 5 October 2026) and its measured cost to complete one index task. Scores are quoted at each model's maximum reasoning effort, which is the tier AA's own model pages display; a lower-effort tier scores and costs much less (bare Kimi K3 is 30.07 at $1.15 per task).

Note on versions. Artificial Analysis rebased this index from v4.1.1 to v4.3.2 during 2026. It is a reset, not a drift — the benchmark basket and the aggregation both changed, so a v4.1.1 score and a v4.3.2 score for the same model on the same day are different numbers, and there is no conversion factor. An earlier version of this page quoted v4.1.1 figures; everything below was re-read against v4.3.2.

ModelInput /1MCached /1MOutput /1MAA Intelligence Index v4.3.2Cost per AA index task
Kimi K3 (max)$3.00$0.30$15.0043.59$2.00
Claude Opus 5 (max)$5.00$0.50$25.0050.78$5.86
GPT-6 Sol (max)$2.00—$10.0047.63$1.04
GPT-5.6 Sol (max) — superseded$4.00$0.40$20.0046.97$1.99
Gemini 3.1 Pro (preview)$2.00 / $4.00 >200k—$12.00 / $18.00 >200k29.72$0.67
Gemini 3.7 Flash (high)$0.75$0.075$3.7539.06$0.93
GLM-5.2 (max)$1.40$0.26$4.4033.71$1.47
GLM-5.3 (max)$1.40$0.26$4.4044.78$2.01
MiMo-V2.6-Pro$0.435—$0.8746.32$0.13
DeepSeek V4-Pro 0813 (off-peak)$0.66$0.022$1.9836.00$0.67
DeepSeek V4-Flash 0731 (off-peak)$0.22$0.007$0.6634.33$0.22

Index and cost-per-task columns read from Artificial Analysis on 5 October 2026 under Intelligence Index v4.3.2, at each model's max reasoning tier — AA's unqualified labels are not a consistent tier across vendors (bare GPT-5.6 Sol is its Low variant at 33.47), so a comparison is only checkable if the tier is named. AA now prices GPT-5.6 Sol at $4/$20, so its $1.99 already reflects OpenAI's 22 August cut; AA has since superseded it with GPT-6 Sol at $2/$10, which is the row to compare against today. Ignore OpenRouter's $2/$10 listing for GPT-5.6 Sol — OpenAI's own docs and AA both say $4/$20. DeepSeek doubles its rates during peak hours, 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday, so its real cost per task is up to 2× the figure shown. Gemini 3.7 Flash pricing is promotional through 31 December 2026 and doubles on 1 January 2027; GPT-5.6 Sol's cut runs at least through 21 November 2026. Cost per task is the figure to compare on — AA runs a different number of tasks per model, so whole-suite totals are not comparable between rows.

Two readings matter, and one of them has reversed since this page first ran.

Against Anthropic, K3 is still the cheap option by a wide margin. It lists 40% below Opus 5 on both sides, and per completed index task the gap is 2.9× — $2.00 against $5.86. Opus 5 leads on the index itself, 50.78 to 43.59, so you are buying seven index points for roughly three times the money. That is a real trade-off, not a free lunch, but the cheap-capable-alternative argument holds.

Against OpenAI's mid-tier it no longer holds. Comparing like with like — max reasoning tier on both sides — GPT-5.6 Sol costs $1.99 per index task against K3's $2.00 and scores higher, 46.97 to 43.59. Under the superseded v4.1.1 numbers K3 looked clearly cheaper; on v4.3.2 that advantage is gone. Part of the reason is token volume: K3 burned 160M output tokens completing the v4.3.2 index against Sol's 90M, so K3's lower per-token rate is cancelled out by writing nearly twice as much.

And GPT-5.6 Sol is already the generous comparison. AA has superseded it with GPT-6 Sol, which at max effort scores 47.63 for $1.04 per index task at $2/$10 list — roughly half K3's cost per task, four index points ahead, and 107 output tokens/second against K3's 45. On price-per-unit-capability K3 is now clearly behind OpenAI's mid-tier, not level with it. If you pick K3 over Sol, pick it for open weights, the 1M-token context or the 3.8-second time to first token (GPT-6 Sol's is 107.76s at max effort) — not for price.

And neither is the cheapest capable option on this table. MiMo-V2.6-Pro scores 46.32 at $0.13 per index task, and GLM 5.3 Flash is in the same territory. DeepSeek V4-Flash is an order of magnitude cheaper again. Frontier-tier list pricing buys you context, licence terms and latency characteristics here — it does not buy you the best score per dollar.

K3's weak spot is throughput. Artificial Analysis's October 2026 measurements clock it at 45.0 output tokens/second at max effort, against 56.1 for Opus 5 and 81.2 for GPT-5.6 Sol. It wins the other half of the latency picture decisively — 3.83s to first token against 59.16s for Opus 5 and 107.55s for Sol at max effort — so K3 feels fast to start and slow to finish. Cheap tokens delivered slowly can still be the wrong trade for an interactive product. For the full capability picture see our Kimi K3 benchmark comparison and the head-to-head on Kimi K3 vs Claude Opus 5.

How do you cut a Kimi K3 bill?

  • Stabilise your prefix. Caching is automatic — there is no cache API to call — but it only fires when the previous request's prompt exceeded 256 tokens and the opening context matches. Any per-request timestamp, request ID, or reordered tool definition at the top of your prompt destroys the hit and costs you 10×. Put volatile content last.
  • Attack output, not input. Output bills at 50× the cached-input rate. Lowering reasoning effort from max to high, or instructing shorter answers, moves the bill more than any context-trimming exercise.
  • Don't plan around a batch discount. K3 is not on the batch endpoint. If you have a large offline job and price sensitivity, K2.7-code at $0.57/$2.40 batched is roughly a quarter of K3's list rate.
  • Meter $web_search. At $0.005 a call plus the tokens the results consume, an agent that searches ten times a turn adds $0.05 per turn before a single token is billed.
  • Compare third parties on more than headline rate. The cheapest OpenRouter listing is 13% below official but ships fp4 with a clipped context. Because K3's native weights are already MXFP4, aggressive quantisation buys little — most providers cluster at list price for a reason.
  • There is no off-peak window. Unlike DeepSeek, Moonshot has no time-of-day discount, so batching work overnight saves nothing.

Is Kimi K3 cheaper to self-host than to buy?

K3 is open-weights under the Kimi K3 License — commercial use is permitted, with two conditions: a Model-as-a-Service business exceeding $20M in revenue over any 12 months must sign a separate agreement with Moonshot, and products above 100 million MAU or $20M monthly revenue must display "Kimi K3" prominently. Internal use is exempt from both.

So self-hosting is legally open. Financially, it is a different question. The Hugging Face checkpoint is 1.56 TB across 96 safetensors shards. At native MXFP4 the weights alone need roughly 1.4 TB of VRAM, and KV cache plus overhead pushes the practical requirement past what an 8×B200 node (1,440 GB) provides — which is why vLLM's published recipe calls for 16.

At specialist-cloud B200 rates near $5.50 per GPU-hour, one always-on replica is 16 × $5.50 = $88/hour, $2,112/day, about $63,400/month. On AWS Capacity Blocks (~$9.36/GPU-hour) the same replica is closer to $108,000. At a 3:1 input-to-output mix of K3's list rates — $6.00 per 1M tokens — $63,400 buys roughly 10.6 billion tokens a month from the API. Heavy prompt caching stretches that figure; nothing about self-hosting does.

The chat product above burns about 1.1 billion tokens a month. You would need roughly 10× that traffic — around a hundred thousand messages a day — before a single self-hosted replica breaks even on raw compute, and that ignores redundancy, a second replica for failover, and the engineers who keep a 16-GPU distributed serving stack alive. Self-hosting K3 wins on data residency, licence certainty, and latency control long before it wins on price.

Can you use Kimi K3 for free?

Yes, but not through the API. The free Adagio tier at kimi.com gives K3 access in the web chat with usage caps — no card required — but excludes Kimi Code, Kimi Claw, and agent swarms. Paid plans are billed per month, cheaper annually:

PlanMonthlyAnnual (per month)What it adds
AdagioFree—K3 chat, file upload, web access; capped agent credits
Moderato$19$152 concurrent agent tasks, 2–4 subagent swarm, Kimi Code
Allegretto$39$31More scheduled tasks, larger swarm
Allegro$99$794 concurrent tasks, 8 subagents, 1M-token extra-long chat, Kimi Claw
Vivace$199$159Highest credit allowance, 25 scheduled tasks

The API and the subscription are separate wallets: a Vivace plan grants you nothing on platform.kimi.ai, and API credit buys you nothing in the app.

So what should you actually pay for?

A decision rule. If your workload is an agent loop over a stable context, buy the API — the automatic 90% cache discount is the single biggest lever on this page and it needs no code. But price the alternatives per completed task, not per token: at $2.00 K3 is a third of Claude Opus 5's cost and worth the seven-point index gap for most work, while GPT-6 Sol beats it on both cost ($1.04 per task) and score (47.63), so K3 has to earn that slot on open weights, the 1M-token window or time-to-first-token rather than price. If the job is a one-shot fan-out over distinct documents, price DeepSeek V4-Pro, GLM-5.2 or MiMo-V2.6-Pro first; you are paying frontier rates for work that rarely needs frontier reasoning, and K3 has no batch discount to soften it. If you are an individual using K3 through the chat app, the free tier is genuinely usable and $19 covers most solo work. And if you are considering self-hosting for cost reasons, run the token math before the hardware math — below roughly 10 billion tokens a month, the API wins outright.

For the architecture, benchmarks, and access options behind these numbers, read the full Kimi K3 guide, or compare it against the other open-weights flagship in DeepSeek V4 vs Kimi K3.

FAQ

How much does Kimi K3 cost?

Kimi K3 costs $3.00 per million input tokens and $15.00 per million output tokens on Moonshot's official API, falling to $0.30 per million on a cache hit. One rate applies across the entire 1,048,576-token context — there is no long-context surcharge and no peak-hour multiplier. Prices exclude applicable tax.

Is Kimi K3 cheaper than Claude Opus 5?

Yes, on both list price and measured cost. K3's $3/$15 undercuts Claude Opus 5's $5/$25 by 40% on each side, and the practical gap is wider still: on Artificial Analysis's Intelligence Index v4.3.2 (read 5 October 2026) one index task costs $2.00 on K3 (max) against $5.86 on Opus 5 (max), about 2.9×. What you give up is capability — Opus 5 leads the index 50.78 to 43.59, a seven-point gap, and generates faster at 56.1 tokens/second against 45.0. The same comparison against OpenAI goes the other way: GPT-5.6 Sol (max) is level on cost at $1.99 per task and ahead on score at 46.97, and its successor GPT-6 Sol (max) is cheaper still at $1.04 for 47.63.

Does Kimi K3 have a free tier?

Not on the API. Moonshot's platform requires a minimum $1 top-up before it will serve requests, with no permanent free allowance and no trial credit; a $5 voucher is granted once cumulative recharges reach $5. The consumer app at kimi.com does have a free Adagio tier with K3 access, file upload and web browsing, subject to usage caps.

Is Kimi K3 free to self-host?

The weights are free to download and the licence permits commercial use, but the compute is not free. The checkpoint is 1.56 TB and needs roughly 16 B200-class GPUs to serve at native MXFP4 — about $88 per hour, or $63,400 a month, at specialist-cloud rates. At a 3:1 input-to-output mix of list prices that same budget buys roughly 10.6 billion API tokens per month, so self-hosting only pays below the API at very high volume — and that comparison ignores redundancy and the staff to run a 16-GPU serving stack.

What is Kimi K3's context window?

Kimi K3 supports 1,048,576 tokens — a full one-million-token context — and Moonshot charges the same per-token rate throughout it. This differs from Gemini 3.1 Pro, which doubles input pricing above 200,000 tokens, and GPT-5.6 Sol, which bills 2× input and 1.5× output on any request exceeding 272,000 input tokens.

How does Kimi K3 cache pricing work?

Caching is automatic — there is no cache API, no manual cache IDs, and no write or storage fee. Moonshot detects a repeated opening context and bills those tokens at $0.30 rather than $3.00. A hit requires the previous request's prompt to have exceeded 256 tokens and the prefix to match byte-for-byte, so any timestamp or request ID near the top of your prompt will silently break it.

Is Kimi K3 available on the batch API?

No. Moonshot's batch endpoint charges 60% of standard pricing but currently supports only K2.5, K2.6 and K2.7-code. If you need a batch discount for high-volume offline work, K2.7-code at $0.57 input and $2.40 output per million is roughly a quarter of K3's real-time list rate.