DeepSeek V4-Flash vs the Frontier: The 2026 Cost Breakdown
Quick answer. DeepSeek V4-Flash, re-released July 31, 2026, costs $0.14 per million input tokens and $0.28 per million output — roughly 10× to 90× cheaper than Western flagships like Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30), while landing near the frontier on coding and agentic benchmarks. Its bigger sibling V4-Pro ($0.435/$0.87) matches GPT-class quality (80.6% SWE-bench Verified) at about 1/34th the price. For most high-volume production workloads, the cost math now clearly favors DeepSeek.
Price alert: DeepSeek rates rise at 16:00 UTC on August 16, 2026. Every DeepSeek price on this page is the current flat rate and holds until then. From August 16 DeepSeek switches to peak/off-peak pricing, with peak hours running 01:00-04:00 and 06:00-10:00 UTC. Off-peak is not a discount on today - it is half of a raised peak, and every tier costs more than the flat rate it replaces.
| Rate (per 1M tokens) | Today | Off-peak from Aug 16 | Peak from Aug 16 |
|---|---|---|---|
| V4-Flash input | $0.14 | $0.22 | $0.44 |
| V4-Flash output | $0.28 | $0.66 | $1.32 |
| V4-Flash cache read | $0.0028 | $0.007 | $0.014 |
| V4-Pro input | $0.435 | $0.66 | $1.32 |
| V4-Pro output | $0.87 | $1.98 | $3.96 |
| V4-Pro cache read | $0.003625 | $0.022 | $0.044 |
Averaged across a 24-hour day (7 peak hours), that is 1.96x today's input, 2.94x output and 7.84x cache reads on V4-Pro, and 2.03x / 3.04x / 3.23x on V4-Flash. Recomputed: V4-Flash goes from 36x cheaper than Claude Opus 5 on input and 89x on output to 23x / 38x off-peak and 11x / 19x at peak (18x / 29x blended). V4-Pro against GPT-5.6 Sol goes from 11.5x / 34.5x to 7.6x / 15.2x off-peak and 3.8x / 7.6x at peak (5.9x / 11.7x blended). Both tables below carry the August 16 rows. Full breakdown in DeepSeek's August 2026 price change.
Something shifted in July 2026. After a month in which OpenAI shipped GPT-5.6, Anthropic launched Claude Opus 5, Moonshot open-weighted Kimi K3, and Google pushed Gemini 3.5, the interesting question stopped being which model is best and became what can you finally afford to build. DeepSeek is the reason. With V4-Pro at general availability and a freshly re-post-trained V4-Flash live as of July 31, DeepSeek is delivering near-frontier capability at prices that undercut every Western lab by an order of magnitude.
This guide compares DeepSeek V4-Flash and V4-Pro head-to-head against the current frontier — on price, on capability, and on the only metric that matters once you're paying real API bills: cost per useful result. Every number below is grounded in vendor pricing and published benchmarks as of August 2026; where a score is the vendor's own harness rather than an independent result, we say so.
What actually shipped: V4-Pro and the July 31 V4-Flash
DeepSeek V4 arrived in two sizes under an MIT license. V4-Pro carries 1.6 trillion total parameters but activates only 49 billion per token; it reached general availability on August 13, 2026, after a preview that opened on April 24. V4-Flash is the smaller, faster sibling — 284 billion total parameters, roughly 13 billion active — and on July 31 DeepSeek re-post-trained it, keeping the same architecture but reporting a significant jump over the earlier Flash preview. Both models ship with a 1M-token context window and automatic prefix caching that makes repeated context nearly free.
The headline isn't the parameter count. It's the price.
How much cheaper is DeepSeek V4, really?
Here is the current API pricing across the models most teams are actually choosing between, per million tokens:
| Model | Input ($/M) | Output ($/M) | Cached input ($/M) |
|---|---|---|---|
| DeepSeek V4-Flash | $0.14 | $0.28 | $0.003 |
| ↳ V4-Flash from Aug 16 (off-peak / peak) | $0.22 / $0.44 | $0.66 / $1.32 | $0.007 / $0.014 |
| DeepSeek V4-Pro | $0.435 | $0.87 | $0.003625 |
| ↳ V4-Pro from Aug 16 (off-peak / peak) | $0.66 / $1.32 | $1.98 / $3.96 | $0.022 / $0.044 |
| Muse Spark 1.2 (contributor) | $0.10 | $0.20 | $0.002 |
| GPT-5.6 Luna | $0.20 | $1.20 | — |
| Gemini 3.5 Flash | $0.75 | $4.50 | — |
| Muse Spark 1.2 (standard) | $1.25 | $4.25 | $0.15 |
| Kimi K3 | $3.00 | $15.00 | $0.30 |
| Claude Opus 5 | $5.00 | $25.00 | — |
| GPT-5.6 Sol | $5.00 | $30.00 | — |
Read that table twice. V4-Flash is 36× cheaper than Claude Opus 5 on input and 89× cheaper on output. V4-Pro — the model that scores 80.6% on SWE-bench Verified — is about 11× cheaper on input and 34× cheaper on output than GPT-5.6 Sol. Only OpenAI's cost-tier Luna (after its July 30 price cut) gets into the same neighborhood, and even Luna costs more on output while sitting a tier below V4-Pro on hard reasoning.
Read the August 16 rows the same way. Once they land, V4-Flash is 23x cheaper than Opus 5 on input and 38x on output off-peak, 11x and 19x at peak (18x and 29x on a 24-hour blended average). V4-Pro against GPT-5.6 Sol drops from 11.5x / 34.5x to 7.6x / 15.2x off-peak and 3.8x / 7.6x at peak. GPT-5.6 Luna is the one rival that takes the lead outright: its $0.20 input undercuts V4-Flash's new $0.22 floor at every hour of the day.
Is it actually good, or just cheap?
Cheap-and-bad is easy. Cheap-and-competitive is the story here. On DeepSeek's own agentic-coding harness at max effort, V4-Flash posts Terminal Bench 2.1 of 82.7, DSBench-FullStack of 68.7, and DeepSWE of 54.4 — numbers that put a $0.14 model in genuine conversation with flagships costing 30× more. V4-Pro goes further: 93.5 on LiveCodeBench, a 3206 Codeforces Elo, 90.1 on GPQA Diamond, and that 80.6% SWE-bench Verified.
These are vendor-published figures produced with DeepSeek's own harness, so treat them as a ceiling rather than a guarantee. But independent developers running V4 in Cursor, Cline, and terminal agents have broadly confirmed the shape: V4-Pro trades blows with GPT-5.6 and Opus 5 on everyday coding, and V4-Flash is startlingly capable for its price. The open question is the top 10–20% of genuinely hard, long-horizon agentic tasks, where the Western flagships still hold an edge.
The independent split verdict on DeepSeek coding. Measured on Vals' neutral SWE-bench Verified harness, the production V4-Pro-0813 checkpoint (GA August 13, 2026) lands #2 of the field at 96.40% ±0.83, behind only Claude Opus 5 at 97.00% and ahead of GPT-5.6 Sol at 96.20% - at $0.022 per test against Opus 5's $1.29. But on LiveBench's agentic-coding column DeepSeek ranks last of seven frontier peers (V4-Pro 54.95, V4-Flash 46.77, against Opus 5's 65.20). Read that as: near-best-in-class on discrete, well-scoped fixes, weakest of its peer group once the task becomes a long autonomous loop. Details in the V4-Pro-0813 guide. One reliability caveat belongs next to those coding numbers: on AA-Omniscience, which measures whether a model knows what it does not know, V4-Pro-0813 scores 0.83 against Claude Opus 5's 37.07 - close to the hallucination floor. Cheap tokens do not buy calibrated confidence.
What does this look like on a real monthly bill?
Abstract multipliers don't pay invoices, so here's a concrete workload: an agentic coding pipeline processing 50M input tokens and 10M output tokens per month.
| Model | Monthly cost (50M in / 10M out) | vs V4-Flash |
|---|---|---|
| DeepSeek V4-Flash | $9.80 | — |
| Muse Spark 1.2 (contributor) | $7.00 | 0.7× |
| ↳ V4-Flash from Aug 16 (off-peak / blended / peak) | $17.60 / $22.73 / $35.20 | 1.8× / 2.3× / 3.6× |
| DeepSeek V4-Pro | $30.45 | 3.1× |
| ↳ V4-Pro from Aug 16 (off-peak / blended / peak) | $52.80 / $68.20 / $105.60 | 5.4× / 7.0× / 10.8× |
| GPT-5.6 Luna | $22.00 | 2.2× |
| Gemini 3.5 Flash | $82.50 | 8.4× |
| Muse Spark 1.2 (standard) | $105.00 | 10.7× |
| Kimi K3 | $300.00 | 31× |
| Claude Opus 5 | $500.00 | 51× |
| GPT-5.6 Sol | $550.00 | 56× |
The same workload that costs $10 on V4-Flash costs $500 on Opus 5. Turn on prefix caching — where repeated system prompts and codebase context hit the $0.003 cached rate — and the DeepSeek bill drops further, toward a blended $0.06 per million. At scale, this is the difference between an AI feature that's economically viable and one that isn't.
After August 16 the same $500 Opus 5 month is 22x a blended V4-Flash month rather than 51x, and 7.3x a blended V4-Pro month rather than 16.4x. The cache line is the one to watch, because it is where the increase is steepest and where DeepSeek's structural edge actually lives: DeepSeek charges 0.83% of its input rate for a cache read on V4-Pro, where every major Western lab charges exactly 10%. On a realistic agentic request (750 fresh input tokens, 290 output, 82,000 read back from cache) V4-Pro costs $0.00088 against Opus 5's $0.052 - 59x, against only 16x on the same work priced statelessly. DeepSeek's advantage roughly quadruples inside an agent loop. From August 16, cache reads rise 7.84x on V4-Pro and 3.2x on V4-Flash, compressing that agentic gap to about 14x (V4-Pro) and 43x (V4-Flash) blended. Worth noting too: third-party OpenRouter V4-Flash endpoints advertising $0.08/$0.18 are actually ~3.4x more expensive per agentic request than DeepSeek's own, because their cache reads cost 5.7x more.
DeepSeek V4-Flash vs V4-Pro: which should you run?
Inside the DeepSeek family the call is simple. Reach for V4-Flash for high-volume, latency-sensitive, well-scoped work: code completion, retrieval-augmented answers, classification, bulk transforms, and most agent steps. Its blended $0.06/M makes it the default for anything you run thousands of times a day. Step up to V4-Pro when a task needs the deepest reasoning — hard SWE-bench-style bug fixes, competition-grade algorithms, or multi-hour agentic runs — and the 3× price difference is worth it for the accuracy. Many teams route the easy 80% of traffic to Flash and escalate the hard 20% to Pro.
Where does Meta's Muse Spark 1.2 sit?
The newest datapoint on this table arrived on August 5, 2026, when Meta shipped Muse Spark 1.2 alongside its Muse Code terminal agent. Its standard tier prices at $1.25 input / $4.25 output per million tokens with cached input at $0.15 and a 3,000 RPM ceiling — roughly 4× under Claude Opus 5 on input and 5.9× under it on output, with no rights to train on your data. That is the part most coverage skipped. The part it didn't skip is the contributor tier at $0.10 / $0.20 (cached $0.002), which goes below V4-Flash on both input and output — the cheapest rate on this page. The discount is paid in kind: Meta trains on your prompts and completions, the rate limit drops to 100 RPM, and the tier is selected by a distinct model id (muse-spark-1.2-contributor) rather than a signed agreement, so it can be switched on with a single config line. Weights are closed today; Meta has announced an open-weights release for Spark 1.2 but has not shipped one. Via OpenRouter the model is meta/muse-spark-1.2, with a 1,048,576-token context window. Our Muse Spark complete guide has the full model reference.
Three caveats travel with those rates, and on a cost page they matter more than the rate card does:
- Cheap tokens are not cheap outcomes. A weaker model burns more turns to reach the same result. Measured on cost per solved task rather than per token, the gap compresses to low single digits — not the 12–21× the published rates imply.
- The capability gap is real. Muse Spark 1.2 loses all three coding benchmarks Meta itself published against Claude Opus 5, and LiveBench scores its agentic coding at 57.6 — its weakest column, and a regression from Spark 1.1's 58.5. That is precisely the capability the agent is marketed on. See Muse Spark 1.2 benchmarks vs Claude Opus 5.
- No subscription and no spend cap. Usage billing only, with no ceiling you can set — unique among the major options here, and a genuine hazard for anyone reading this page to control a bill.
Where it genuinely wins is efficiency per result rather than raw rate: Muse Spark 1.2 now ranks 5th on the Vals Index at 71.88% and carries the lowest cost per test in the top five, about $0.69. That is a real credential, and it is the honest reason to keep it on the shortlist next to DeepSeek V4.
When should you still pay for a Western flagship?
Cost isn't the only axis, and being honest about the trade-offs is the point:
- Multimodality. DeepSeek V4 is text-in, text-out. If you need native image, audio, or video input, Claude Opus 5, GPT-5.6, and Kimi K3 (which ships native vision) are the picks.
- The hardest long-horizon agents. For agents that must run autonomously for hours or days on genuinely novel problems, Opus 5 and GPT-5.6 Sol still lead. See our Claude Opus 5 launch guide for where that tier earns its price.
- Data residency and compliance. DeepSeek's API is hosted in China; regulated workloads may require self-hosting the open weights or choosing a Western provider. The good news: V4's MIT license means you can self-host.
- Ecosystem lock-in. If your stack is deep in the OpenAI or Anthropic SDK, switching cost is real — though DeepSeek's API is OpenAI-compatible, which softens the migration.
How does DeepSeek V4 compare to each rival, one on one?
We've broken out the head-to-head matchups that matter most for cost-conscious developers:
- DeepSeek V4-Flash vs Claude Opus 5 — the widest cost gap in AI right now.
- DeepSeek V4 vs GPT-5.6 (Sol, Terra, Luna) — how the tiers stack up on price and coding.
- DeepSeek V4 vs Kimi K3 — two open-weight giants, very different price points.
- DeepSeek V4 vs Gemini 3.5 — cost-per-task for everyday developer workloads.
For the full model reference — architecture, self-hosting, hardware, and IDE setup — see our continuously updated DeepSeek V4 complete guide.
The bottom line
DeepSeek V4 hasn't dethroned the frontier on raw capability — Opus 5 and GPT-5.6 Sol still win the hardest tasks. What it has done is collapse the price of good enough to a rounding error. When V4-Flash delivers 80%+ of flagship quality at 2–3% of the cost, the default for high-volume production work flips. The right architecture for most teams in late 2026 is DeepSeek V4 for the bulk of traffic, with a frontier model on standby for the cases that genuinely need it.
Want the full picture? Read our continuously-updated DeepSeek V4 complete guide — architecture, benchmarks, self-hosting, hardware requirements, and IDE integration for both V4-Pro and V4-Flash.
FAQ
How much does DeepSeek V4-Flash cost?
V4-Flash is $0.14 per million input tokens and $0.28 per million output tokens, with cached input at about $0.003 per million (a 98% discount). That works out to a blended rate near $0.06 per million on cache-heavy workloads — roughly an order of magnitude below every Western flagship. From 16:00 UTC on August 16, 2026 those rates rise to $0.22 input / $0.66 output / $0.007 cached off-peak and $0.44 / $1.32 / $0.014 at peak (peak hours are 01:00–04:00 and 06:00–10:00 UTC).
Is DeepSeek V4 cheaper than GPT-5.6 and Claude Opus 5?
Dramatically. V4-Flash is about 36× cheaper than Opus 5 on input and 89× cheaper on output. V4-Pro, which scores 80.6% on SWE-bench Verified, is roughly 11× cheaper on input and 34× cheaper on output than GPT-5.6 Sol. Only OpenAI's cost-tier Luna comes close, and it still costs more on output. From August 16, 2026 those gaps narrow: V4-Flash becomes 23x / 38x cheaper than Opus 5 off-peak and 11x / 19x at peak, and V4-Pro becomes 7.6x / 15.2x cheaper than Sol off-peak and 3.8x / 7.6x at peak.
Is DeepSeek V4 actually as good as the frontier models?
Close, not identical. On vendor-published coding and agentic benchmarks V4-Pro trades blows with GPT-5.6 and Opus 5, and V4-Flash is remarkably capable for its price. The gap remains on the hardest long-horizon agentic tasks and on multimodality — DeepSeek V4 is text-only, while the Western flagships and Kimi K3 handle images.
Is Meta's Muse Spark 1.2 cheaper than DeepSeek V4-Flash?
Only on the contributor tier. At $0.10/$0.20 per million tokens it undercuts V4-Flash's $0.14/$0.28, but Meta trains on your prompts and completions and caps you at 100 RPM. The standard tier — $1.25/$4.25, no training on your data — is roughly 9× V4-Flash's input rate and 15× its output rate, though it still undercuts Claude Opus 5 by 5.9× on output.
What's the difference between V4-Pro and V4-Flash?
V4-Pro is the 1.6T-parameter reasoning model (49B active) for the hardest tasks; V4-Flash is the 284B model (13B active) tuned for speed and volume at roughly a third of Pro's price. Route routine, high-frequency work to Flash and escalate genuinely hard problems to Pro.
Can I self-host DeepSeek V4?
Yes. V4 ships under an MIT license with open weights, so you can run it on your own hardware for data-residency or compliance reasons. See the complete guide for VRAM and GPU requirements.
Building cost-efficient AI features but short on senior engineers?
Codersera matches you with vetted, remote-ready developers who ship with models like DeepSeek V4, Claude, and GPT in their daily workflow. Extend your team in days, not months — risk-free.