Kimi K3 Pricing: API Costs, Plans & Real Bills (2026)
What does Kimi K3 cost per million tokens?
Moonshot AI publishes a single rate card for Kimi K3. There is no "mini", no "turbo", and no separate thinking-mode SKU — K3's reasoning is always on and its thinking tokens bill as ordinary output tokens. Reasoning depth is a request parameter (low, high, max), not a price tier.
| Access route | Input /1M | Cached input /1M | Output /1M | Context | Notes |
|---|---|---|---|---|---|
Kimi K3 — official API (kimi-k3) | $3.00 | $0.30 | $15.00 | 1,048,576 | Flat across the whole window; prices exclude tax |
| OpenRouter — cheapest listing | $2.60 | $0.29 | $13.00 | ~1M | Sail Research; fp4, slightly clipped context |
| OpenRouter — typical provider | $3.00 | $0.29 | $15.00 | 1,048,576 | Fireworks, Together, DeepInfra, Baseten, Moonshot's own endpoint |
| OpenRouter — most expensive | $6.00 | — | $22.50 | ~1M | Morph Fast; latency-optimised |
| Batch API | Not available for K3 — Moonshot's 60%-of-list batch tier covers only K2.5, K2.6 and K2.7-code | ||||
$web_search built-in tool | $0.005 per successful call, plus the returned results billed as normal input tokens | ||||
Rates in USD per 1M tokens, verified against Moonshot's K3 pricing page and OpenRouter on 23 August 2026.
Two things are unusual here. First, the cache-hit rate is a straight 90% discount with no cache-write premium and no storage fee — unlike Anthropic, which charges 1.25× to write a cache entry. Second, Moonshot does not tier by context length. Many Chinese labs do, and DeepSeek now runs a 2× peak-hour multiplier; K3 does neither. The $3 you pay on a 2,000-token prompt is the $3 you pay on a 900,000-token one.
What do you have to pay before the API will answer?
There are no free API credits. You top up a balance, and your recharge total determines your rate limits. The minimum is $1, and cumulative recharges reaching $5 earn a $5 voucher (vouchers themselves do not count toward the tier threshold).
| Tier | Cumulative recharge | Concurrency | RPM | TPM | Tokens/day |
|---|---|---|---|---|---|
| Tier 0 | $1 | 1 | 3 | 500,000 | 1,500,000 |
| Tier 1 | $10 | 50 | 200 | 2,000,000 | Unlimited |
| Tier 2 | $20 | 100 | 500 | 3,000,000 | Unlimited |
| Tier 3 | $100 | 200 | 5,000 | 3,000,000 | Unlimited |
| Tier 4 | $1,000 | 400 | 5,000 | 4,000,000 | Unlimited |
| Tier 5 | $3,000 | 1,000 | 10,000 | 5,000,000 | Unlimited |
The step that matters is $1 → $10. Tier 0 gives you 3 requests per minute and a 1.5M-token daily ceiling, which is a smoke-test allowance, not a workload. Ten dollars buys 50-way concurrency and removes the daily cap entirely. If you are evaluating K3 for an agent loop, put in $10 before you conclude it is slow.
What does Kimi K3 actually cost on real workloads?
List price tells you very little until you multiply it by a token profile. Three worked examples, arithmetic shown.
A coding-agent session over a repository
Assume 40 turns, each resending a growing context averaging 60,000 prompt tokens, with 1,500 output tokens per turn. That is 2.4M input and 60k output.
- No caching: 2.4 × $3.00 = $7.20, plus 0.06 × $15.00 = $0.90 → $8.10
- At a 90% cache-hit rate (Moonshot's stated figure for coding traffic on its Mooncake serving stack): 2.16 × $0.30 = $0.65, plus 0.24 × $3.00 = $0.72, plus $0.90 output → $2.27
Caching is worth 3.6× here because the prefix is enormous and stable. This is the workload K3's pricing is shaped for.
A document-analysis batch job
1,000 documents, ~50,000 tokens each, a shared 2,000-token system prompt, 800 tokens of output apiece.
- Cached prefix: 999 × 2,000 = 2.0M at $0.30 → $0.60
- Uncached document bodies: 50.0M at $3.00 → $150.00
- Output: 0.8M at $15.00 → $12.00
- Total ≈ $162.60, or $0.163 per document
Notice how little the cache saved: 0.4%. Caching rewards repeated prefixes, and a fan-out over distinct documents has almost none. Worse, K3 is excluded from Moonshot's batch tier, so there is no 40% discount to fall back on — a job you could halve on OpenAI or Anthropic's batch endpoints, you pay full freight for here.
A chat product at 10,000 messages a day
Assume 3,200 prompt tokens per message (system prompt plus a few turns of history) and 350 output tokens, with a realistic 60% cache-hit rate.
- Cached input: 19.2M at $0.30 → $5.76
- Uncached input: 12.8M at $3.00 → $38.40
- Output: 3.5M at $15.00 → $52.50
- $96.66/day ≈ $2,900/month, or $0.0097 per message
Output is 54% of that bill on 10% of the tokens. For chat, the lever is response length and reasoning effort — not context trimming.
Is Kimi K3 cheaper than Claude Opus 5, GPT-5.6 and Gemini?
On list price K3 sits mid-pack among frontier models and far above the discount Chinese tier. Cheaper-is-better is the wrong reading, so the last two columns give a capability-adjusted view using Artificial Analysis's Intelligence Index and its measured cost to complete one benchmark task.
| Model | Input /1M | Cached /1M | Output /1M | AA Intelligence Index | Cost per AA task |
|---|---|---|---|---|---|
| Kimi K3 | $3.00 | $0.30 | $15.00 | 60 | $0.84 |
| Claude Opus 5 | $5.00 | $0.50 | $25.00 | 63 | $2.34 |
| GPT-5.6 Sol | $4.00 | $0.40 | $20.00 | 61 | $1.23 * |
| Gemini 3.1 Pro | $2.00 / $4.00 >200k | — | $12.00 / $18.00 >200k | — | — |
| Gemini 3.7 Flash | $0.75 | $0.075 | $3.75 | — | — |
| GLM-5.2 | $1.40 | $0.26 | $4.40 | — | — |
| DeepSeek V4-Pro (off-peak) | $0.66 | $0.022 | $1.98 | — | — |
| DeepSeek V4-Flash (off-peak) | $0.22 | $0.007 | $0.66 | — | — |
* Artificial Analysis measured GPT-5.6 Sol before OpenAI's 22 August price cut (its blended $4.35/1M implies the old $5/$30 rates); at the new $4/$20 the cost per task should fall to roughly $0.87. DeepSeek doubles its rates during peak hours, 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday. Gemini 3.7 Flash pricing is promotional through 31 December 2026 and doubles on 1 January 2027; GPT-5.6 Sol's cut runs at least through 21 November 2026.
The pattern is the useful part. K3 lists 40% below Opus 5 on both sides, but the per-task gap is 2.8× — because K3 answers more tersely, so a given task consumes fewer billed tokens. Against GPT-5.6 Sol the picture is much closer once the August price cut is applied: K3's advantage narrows from roughly 1.5× to near parity. And nothing here touches DeepSeek V4-Flash, which is an order of magnitude cheaper and a different class of model.
K3's weak spot is throughput. Artificial Analysis clocks it at 37.2 output tokens/second against 55.5 for Opus 5 and 71.8 for GPT-5.6 Sol. Cheap tokens delivered slowly can still be the wrong trade for an interactive product. For the full capability picture see our Kimi K3 benchmark comparison and the head-to-head on Kimi K3 vs Claude Opus 5.
How do you cut a Kimi K3 bill?
- Stabilise your prefix. Caching is automatic — there is no cache API to call — but it only fires when the previous request's prompt exceeded 256 tokens and the opening context matches. Any per-request timestamp, request ID, or reordered tool definition at the top of your prompt destroys the hit and costs you 10×. Put volatile content last.
- Attack output, not input. Output bills at 50× the cached-input rate. Lowering reasoning effort from
maxtohigh, or instructing shorter answers, moves the bill more than any context-trimming exercise. - Don't plan around a batch discount. K3 is not on the batch endpoint. If you have a large offline job and price sensitivity, K2.7-code at $0.57/$2.40 batched is roughly a quarter of K3's list rate.
- Meter
$web_search. At $0.005 a call plus the tokens the results consume, an agent that searches ten times a turn adds $0.05 per turn before a single token is billed. - Compare third parties on more than headline rate. The cheapest OpenRouter listing is 13% below official but ships fp4 with a clipped context. Because K3's native weights are already MXFP4, aggressive quantisation buys little — most providers cluster at list price for a reason.
- There is no off-peak window. Unlike DeepSeek, Moonshot has no time-of-day discount, so batching work overnight saves nothing.
Is Kimi K3 cheaper to self-host than to buy?
K3 is open-weights under the Kimi K3 License — commercial use is permitted, with two conditions: a Model-as-a-Service business exceeding $20M in revenue over any 12 months must sign a separate agreement with Moonshot, and products above 100 million MAU or $20M monthly revenue must display "Kimi K3" prominently. Internal use is exempt from both.
So self-hosting is legally open. Financially, it is a different question. The Hugging Face checkpoint is 1.56 TB across 96 safetensors shards. At native MXFP4 the weights alone need roughly 1.4 TB of VRAM, and KV cache plus overhead pushes the practical requirement past what an 8×B200 node (1,440 GB) provides — which is why vLLM's published recipe calls for 16.
At specialist-cloud B200 rates near $5.50 per GPU-hour, one always-on replica is 16 × $5.50 = $88/hour, $2,112/day, about $63,400/month. On AWS Capacity Blocks (~$9.36/GPU-hour) the same replica is closer to $108,000. Against Artificial Analysis's blended API rate of $2.31 per 1M tokens, $63,400 buys roughly 27 billion tokens a month from the API.
The chat product above burns about 1.1 billion tokens a month. You would need roughly 25× that traffic — around a quarter of a million messages a day — before a single self-hosted replica breaks even on raw compute, and that ignores redundancy, a second replica for failover, and the engineers who keep a 16-GPU distributed serving stack alive. Self-hosting K3 wins on data residency, licence certainty, and latency control long before it wins on price.
Can you use Kimi K3 for free?
Yes, but not through the API. The free Adagio tier at kimi.com gives K3 access in the web chat with usage caps — no card required — but excludes Kimi Code, Kimi Claw, and agent swarms. Paid plans are billed per month, cheaper annually:
| Plan | Monthly | Annual (per month) | What it adds |
|---|---|---|---|
| Adagio | Free | — | K3 chat, file upload, web access; capped agent credits |
| Moderato | $19 | $15 | 2 concurrent agent tasks, 2–4 subagent swarm, Kimi Code |
| Allegretto | $39 | $31 | More scheduled tasks, larger swarm |
| Allegro | $99 | $79 | 4 concurrent tasks, 8 subagents, 1M-token extra-long chat, Kimi Claw |
| Vivace | $199 | $159 | Highest credit allowance, 25 scheduled tasks |
The API and the subscription are separate wallets: a Vivace plan grants you nothing on platform.kimi.ai, and API credit buys you nothing in the app.
So what should you actually pay for?
A decision rule. If your workload is an agent loop over a stable context, buy the API — the 90% cache discount and K3's token efficiency make it the cheapest capable option in the frontier tier, and the $0.84-per-task figure is real. If it is a one-shot fan-out over distinct documents, price DeepSeek V4-Pro or GLM-5.2 first; you are paying frontier rates for a job that rarely needs frontier reasoning, and K3 has no batch discount to soften it. If you are an individual using K3 through the chat app, the free tier is genuinely usable and $19 covers most solo work. And if you are considering self-hosting for cost reasons, run the token math before the hardware math — below roughly 27 billion tokens a month, the API wins outright.
For the architecture, benchmarks, and access options behind these numbers, read the full Kimi K3 guide, or compare it against the other open-weights flagship in DeepSeek V4 vs Kimi K3.
FAQ
How much does Kimi K3 cost?
Kimi K3 costs $3.00 per million input tokens and $15.00 per million output tokens on Moonshot's official API, falling to $0.30 per million on a cache hit. One rate applies across the entire 1,048,576-token context — there is no long-context surcharge and no peak-hour multiplier. Prices exclude applicable tax.
Is Kimi K3 cheaper than Claude Opus 5?
Yes, on both list price and measured cost. K3's $3/$15 undercuts Claude Opus 5's $5/$25 by 40% on each side. The practical gap is wider: Artificial Analysis measures $0.84 per benchmark task for K3 against $2.34 for Opus 5, because K3 produces fewer billed tokens per answer. Opus 5 scores slightly higher on intelligence (63 vs 60) and generates faster.
Does Kimi K3 have a free tier?
Not on the API. Moonshot's platform requires a minimum $1 top-up before it will serve requests, with no permanent free allowance and no trial credit; a $5 voucher is granted once cumulative recharges reach $5. The consumer app at kimi.com does have a free Adagio tier with K3 access, file upload and web browsing, subject to usage caps.
Is Kimi K3 free to self-host?
The weights are free to download and the licence permits commercial use, but the compute is not free. The checkpoint is 1.56 TB and needs roughly 16 B200-class GPUs to serve at native MXFP4 — about $88 per hour, or $63,400 a month, at specialist-cloud rates. That equals roughly 27 billion API tokens per month, so self-hosting only pays below the API at very high volume.
What is Kimi K3's context window?
Kimi K3 supports 1,048,576 tokens — a full one-million-token context — and Moonshot charges the same per-token rate throughout it. This differs from Gemini 3.1 Pro, which doubles input pricing above 200,000 tokens, and GPT-5.6 Sol, which bills 2× input and 1.5× output on any request exceeding 272,000 input tokens.
How does Kimi K3 cache pricing work?
Caching is automatic — there is no cache API, no manual cache IDs, and no write or storage fee. Moonshot detects a repeated opening context and bills those tokens at $0.30 rather than $3.00. A hit requires the previous request's prompt to have exceeded 256 tokens and the prefix to match byte-for-byte, so any timestamp or request ID near the top of your prompt will silently break it.
Is Kimi K3 available on the batch API?
No. Moonshot's batch endpoint charges 60% of standard pricing but currently supports only K2.5, K2.6 and K2.7-code. If you need a batch discount for high-volume offline work, K2.7-code at $0.57 input and $2.40 output per million is roughly a quarter of K3's real-time list rate.