DeepSeek V4 vs Kimi K3: Two Open Giants (2026)

Quick answer. DeepSeek V4-Pro now costs $0.66 input / $1.98 output per million tokens off-peak ($1.32 / $3.96 peak) versus Kimi K3's flat $3 / $15 — roughly 3.5x cheaper on input and 5.9x cheaper on output on a 24-hour blend (as of August 20, 2026). Pick K3 when you need native vision or its best-in-class front-end coding; pick V4 for text-only work where cost dominates.

Price change: DeepSeek rates rose at 16:00 UTC on August 16, 2026. DeepSeek now runs peak/off-peak pricing, with peak hours 01:00-04:00 and 06:00-10:00 UTC. Off-peak is not a discount on the old flat rate - it is half of a raised peak, and every tier costs more than the flat rate it replaced. All DeepSeek prices on this page use the current peak/off-peak rates.

Rate (per 1M tokens)Flat rate (until Aug 16)Off-peak (current)Peak (current)
V4-Flash input$0.14$0.22$0.44
V4-Flash output$0.28$0.66$1.32
V4-Flash cache read$0.0028$0.007$0.014
V4-Pro input$0.435$0.66$1.32
V4-Pro output$0.87$1.98$3.96
V4-Pro cache read$0.003625$0.022$0.044

Averaged across a 24-hour day (7 peak hours), the new rates work out to 1.96x the old flat input, 2.94x output and 7.84x cache reads on V4-Pro, and 2.03x / 3.04x / 3.23x on V4-Flash. Against Kimi K3: V4-Pro's old 6.9x input / 17.2x output edge is now 4.5x / 7.6x off-peak and 2.3x / 3.8x at peak (3.5x / 5.9x blended), and the worked monthly bill below moved from $30.45 to $52.80 off-peak, $105.60 at peak, $68.20 blended - so K3 now costs about 4.4x V4-Pro rather than about 10x. The verdict holds; the margin more than halved. Full breakdown in DeepSeek's August 2026 price change.

Two of 2026's most talked-about open-weight models come from the same corner of the world and land at wildly different price points. Kimi K3 (Moonshot AI) is a 2.8-trillion-parameter mixture-of-experts model with native vision that topped the blind Frontend Code Arena, beating leading US models. DeepSeek V4 ships in two text-only tiers — V4-Pro and the even cheaper V4-Flash — at a fraction of K3's per-token cost.

Both have open weights, so both are self-hostable. The real decision isn't "which is better in the abstract" — it's when Kimi K3's vision and front-end coding earn their premium, and when DeepSeek V4's price makes it the obvious default. This is a cost-efficiency comparison: we lead with price per task, then stay honest about where the pricier model wins.

How much cheaper is DeepSeek V4 than Kimi K3?

Here are the current list prices (as of August 20, 2026), per million tokens, side by side. DeepSeek's rates are peak/off-peak; Kimi K3's are flat around the clock.

ModelInputOutputCached inputVisionOpen weights
DeepSeek V4-Flash$0.22 off-peak / $0.44 peak$0.66 / $1.32$0.007 / $0.014NoYes
DeepSeek V4-Pro$0.66 off-peak / $1.32 peak$1.98 / $3.96$0.022 / $0.044NoYes (MIT)
Kimi K3$3.00 (flat)$15.00$0.30NativeYes

The multipliers are still substantial, if narrower than before August 16. Against Kimi K3, DeepSeek V4-Pro is about 4.5x cheaper on input and 7.6x cheaper on output off-peak ($0.66 / $1.98 vs $3 / $15), 2.3x / 3.8x at peak — call it 3.5x / 5.9x on a 24-hour blend. Drop to V4-Flash and the gap widens to roughly 14x on input and 23x on output off-peak (7x / 11x at peak). If your workload is pure text — code generation, refactoring, RAG, document processing, agent tool-calling — that difference compounds fast.

Kimi K3's cached-input rate of $0.30 helps on repeated-context workloads, but it's still roughly 11x DeepSeek's blended cache rate (about $0.028; $0.022 off-peak, $0.044 peak) — 13.6x off-peak, 6.8x at peak. DeepSeek's cache discount (reads at ~3.3% of the input rate, a third of the Western 10% norm) remains one of the most aggressive in the market and is the single biggest lever for cutting a real monthly bill.

Cache is where the August 16 change bit hardest. Before the change DeepSeek priced a cache read at just 0.83% of its input rate on V4-Pro, where every major Western lab prices cache reads at exactly 10% of input. That was why DeepSeek's advantage roughly quadrupled inside an agent loop: on a realistic agentic request (750 fresh input tokens, 290 output, 82,000 read back from cache) V4-Pro cost $0.00088 against Claude Opus 5's $0.052 - a 59x gap, versus only 16x on the same work priced statelessly. Since August 16 the cache read is 7.84x higher at a blended $0.0284, so Kimi K3's $0.30 fell from 83x V4-Pro's cache rate to about 11x (13.6x off-peak, 6.8x at peak), and the agentic gap against Opus 5 compressed from 59x to roughly 14x. Cache reads still cost only ~3.3% of input - a third of the Western norm - but the structural edge is a fraction of what it was.

What does a real monthly bill look like?

Numbers per million tokens are abstract. Here's a concrete workload — 50M input tokens + 10M output tokens per month, a realistic footprint for a small product team running an AI coding assistant or an internal agent — priced across the field.

ModelMonthly cost (50M in / 10M out)vs V4-Pro
DeepSeek V4-Flash$22.72 blended ($17.60 off-peak)0.3x
DeepSeek V4-Pro$68.20 blended ($52.80 off-peak / $105.60 peak)1x (baseline)
GPT-5.6 Luna$22.000.3x
Gemini 3.5 Flash$82.501.2x
Kimi K3$300.00~4.4x
Claude Opus 5$500.00~7.3x
GPT-5.6 Sol$550.00~8.1x

Kimi K3 lands at $300/month on this workload — roughly 4.4x the blended cost of DeepSeek V4-Pro and about 13x the cost of V4-Flash. It's still meaningfully cheaper than Claude Opus 5 ($500) or GPT-5.6 Sol ($550 at the official $5 / $30 list rate; OpenRouter currently lists Sol at $2.50 / $15 — OpenAI's discounted Flex-tier rate — which would put Sol near $275 via that route), which is exactly K3's positioning: a frontier-tier open model that undercuts the Western flagships while offering something they and DeepSeek don't — native vision.

For a deeper cost breakdown of DeepSeek V4 against every Western flagship, the DeepSeek V4-Flash cost-vs-frontier guide is our flagship cost hub, and the complete DeepSeek V4 guide covers the architecture end to end.

When is Kimi K3 worth the premium?

Paying 4-5x more only makes sense when you're buying capability DeepSeek V4 physically cannot provide. Two cases stand out.

1. You need native vision

This is the clean-cut case. DeepSeek V4 — both Pro and Flash — is text-only. No image, audio, or video input. If your product processes screenshots, design mockups, PDFs-as-images, charts, UI states, or any multimodal input, DeepSeek V4 is off the table regardless of price. Kimi K3 has native vision built in, and so do Claude Opus 5 and the GPT-5.6 family. Among the vision-capable options, K3 is the cheapest, which reframes its $3/$15 pricing as a bargain rather than a premium — if vision is a hard requirement.

2. You're shipping front-end code and want the current arena leader

Kimi K3 ranked #1 on the blind Frontend Code Arena at 1,679 Elo, ahead of Claude Fable 5, in developer blind testing. For teams whose primary workload is generating polished React/Next.js components, layouts, and UI, that measured edge can justify the spend — especially if you're pairing it with vision to feed the model a design mock and get matching markup back.

That said, DeepSeek V4-Pro is no slouch on code: LiveCodeBench 93.5, Codeforces Elo 3206, SWE-bench Verified 80.6. For back-end logic, algorithms, and repository-scale engineering tasks, V4-Pro is competitive at a fraction of the cost. The K3 premium is most defensible narrowly on the front-end/visual slice, not coding across the board. We break the coding-specific tradeoffs down further in Kimi K2.6 vs DeepSeek V4 vs GLM 5.1.

The independent split verdict on DeepSeek coding. Those are DeepSeek's own harness figures. Measured on Vals' neutral SWE-bench Verified harness, the production V4-Pro-0813 checkpoint (GA August 13, 2026) lands #2 of the field at 96.40% ±0.83, behind only Claude Opus 5 at 97.00% and ahead of GPT-5.6 Sol at 96.20% - at $0.022 per test against Opus 5's $1.29. But on LiveBench's agentic-coding column DeepSeek ranks last of seven frontier peers (V4-Pro 54.95, V4-Flash 46.77, against Opus 5's 65.20). Read that as: near-best-in-class on discrete, well-scoped fixes, weakest of its peer group once the task becomes a long autonomous loop. Details in the V4-Pro-0813 guide.

When does DeepSeek V4 win outright?

For the large majority of text workloads, DeepSeek V4 is the rational default:

  • Cost-sensitive scale. High-volume text generation, classification, summarization, RAG, and agent orchestration where per-token cost is the constraint. At about $68/month blended vs $300 for the same job, V4-Pro frees budget for everything else.
  • Cache-heavy pipelines. If you're re-sending the same system prompt or document context repeatedly, DeepSeek's cached input ($0.022 off-peak / $0.044 peak) still pulls the effective blended rate far below anything else in this table.
  • Back-end and repo-scale engineering. V4-Pro's SWE-bench Verified 80.6 and Codeforces 3206 make it a strong coding model for anything that isn't specifically front-end/visual.
  • Even tighter budgets. V4-Flash at $0.22/$0.66 off-peak takes the same monthly job to about $23 blended (under $18 if you keep traffic off-peak). It's a re-post-trained release (July 31, 2026) with a significant reported jump on DeepSeek's own harness — Terminal Bench 2.1 at 82.7 and DSBench-FullStack at 68.7.

DeepSeek V4-Pro also carries an MIT license with open weights, so if data residency or compliance rules out the hosted API (which runs in China), you can self-host. That's a genuine escape hatch Western flagships don't offer. See the V4-Pro permanent price cut for how that pricing became the standard list rate rather than a temporary discount.

Can I self-host either model?

Both models have open weights, so yes — but the hardware bill is where they diverge again.

Kimi K3 is a 2.8T-parameter MoE that activates about 1.8% of its experts per token. Moonshot shipped it with quantization-aware training (MXFP4 weights, MXFP8 activations) specifically for broad hardware compatibility, which softens the deployment burden — but a 2.8T model still demands a serious multi-GPU cluster to serve at usable latency. DeepSeek V4-Pro (1.6T total, 49B active) and especially V4-Flash (284B total, ~13B active) are markedly lighter to run. For most teams, DeepSeek V4-Flash is the realistic self-host target; K3 self-hosting is a data-center project.

Either way, the hosted APIs are cheap enough that self-hosting is usually a compliance or scale decision, not a cost one. DeepSeek's API is also OpenAI-compatible, which softens migration if you're moving off another provider.

The bottom line on cost per task

July 2026 shifted the question from "which model is best" to "what can you finally afford to build." DeepSeek V4 is why: near-frontier text capability at a large multiple below Western flagship pricing even after its August 16 rate rise, and roughly 3.5-6x cheaper than Kimi K3 on the same tokens (blended; 4.5-7.6x off-peak). If your workload is text — which most are — V4-Pro or V4-Flash is the default, and the savings fund the rest of your stack.

Kimi K3 earns its premium in exactly two lanes: when you need native vision, and when front-end/visual code quality is the whole point. In both, it's the cheapest capable option — cheaper than Opus 5 or GPT-5.6 Sol for the same multimodal frontier tier. Buy K3 for what it can do that V4 can't, not as a general upgrade.

Choosing and wiring the right model into a product is the easy part once you know your cost envelope. Getting an engineering team that can actually ship it — build the pipelines, tune the prompts, handle self-hosting and evals — is the harder problem. Hire vetted remote developers through Codersera to extend your team with engineers who know how to put these models into production.

FAQ

Is DeepSeek V4 cheaper than Kimi K3?

Yes, substantially — though the gap narrowed on August 16, 2026, when DeepSeek moved to peak/off-peak pricing. DeepSeek V4-Pro now costs $0.66 / $1.98 off-peak ($1.32 / $3.96 peak) per million tokens versus Kimi K3's flat $3 / $15 — about 3.5x cheaper on input and 5.9x on output on a 24-hour blend. DeepSeek V4-Flash ($0.22 / $0.66 off-peak) is cheaper still. On a 50M-input / 10M-output monthly workload, V4-Pro costs about $68 blended against roughly $300 for Kimi K3.

Does DeepSeek V4 have vision like Kimi K3?

No. Both DeepSeek V4-Pro and V4-Flash are text-only — no image, audio, or video input. Kimi K3 has native vision. If your workload requires multimodal input, DeepSeek V4 is not an option and Kimi K3 (or Claude Opus 5 / GPT-5.6) is the pick. Among vision-capable models, K3 is the cheapest.

Which is better for coding, DeepSeek V4 or Kimi K3?

It depends on the coding task. Kimi K3 ranked #1 on the blind Frontend Code Arena (1,679 Elo, ahead of Claude Fable 5), so it leads on front-end/visual code. DeepSeek V4-Pro is strong across the board — LiveCodeBench 93.5, SWE-bench Verified 80.6, Codeforces Elo 3206 — at roughly a tenth of the cost, making it the better value for back-end and repo-scale engineering.

Are both models open-weight and self-hostable?

Yes. Kimi K3's 2.8T weights were released on July 27, 2026, and DeepSeek V4-Pro ships under an MIT license with open weights. Both can be self-hosted, but K3's 2.8T size makes it a data-center-scale deployment, while DeepSeek V4-Flash (284B total, ~13B active) is far lighter and the realistic self-host target for most teams.

Why is Kimi K3 still cheaper than Claude Opus 5?

Kimi K3 at $3 / $15 undercuts Claude Opus 5 ($5 / $25) and GPT-5.6 Sol ($5 / $30 official list; $2.50 / $15 via OpenRouter, matching OpenAI's Flex tier) while matching the multimodal frontier tier. On a 50M / 10M monthly workload K3 costs about $300 versus $500 for Opus 5 and $550 for Sol at list rates. K3's premium is only steep relative to DeepSeek V4 — against Western flagships it's the value option in the vision-capable frontier bracket.