DeepSeek V4 vs GPT-5.6: Real Cost Per 1M Tokens (2026)
Quick answer. DeepSeek V4-Pro costs $0.66 input / $1.98 output per 1M tokens off-peak on DeepSeek's own API — roughly 6x cheaper on input and 10x on output than GPT-5.6 Sol's $4/$20. Peak hours (01:00-04:00 and 06:00-10:00 UTC, Monday to Friday) double those rates, so 79% of the week is off-peak.
What changed since the last update to this page (3 September 2026):
- OpenAI launched GPT-6 Astra today, above the whole GPT-5.6 family, at $10 input / $50 output per 1M tokens (short context) and $20 / $75 above 272K input tokens. It is added to every cost table below. Full detail in the GPT-6 Astra complete guide.
- GPT-5.6 Sol's promotional pricing still stands. OpenAI states it is available at least through 21 November 2026. Sol's live rate is $4 / $20 short context, not the $5 / $30 list price this page previously used.
- The peak/off-peak blend on this page has been corrected. DeepSeek's peak window runs Monday to Friday only, so 79% of the week is off-peak — not the ~71% a naive 24-hour average implies.
- Every price below is vendor-direct, taken from DeepSeek's, OpenAI's and Anthropic's own pricing pages on 3 September 2026. See the first section for why that matters more than it sounds.
Cost comparisons between DeepSeek and the Western labs get repeated a lot and verified almost never. The numbers move — DeepSeek repriced in August, OpenAI has a promotional rate with an expiry date, and a new top-of-market model landed today — and the figures that circulate most widely are frequently not the ones the vendor actually bills.
This page compares DeepSeek's two V4 tiers against all three GPT-5.6 tiers and GPT-6 Astra on the one axis that decides your budget: what you actually pay per million tokens, on each vendor's own API, at each vendor's own rates.
Which price are you actually reading?
Start here, because it invalidates a lot of the comparisons you will find elsewhere.
DeepSeek V4-Flash is widely listed at around $0.089 input / $0.177 output per 1M tokens. Those numbers appear on aggregators, in model-comparison tables, and in a good deal of coverage. They are not what DeepSeek's own API charges. DeepSeek's published rate for V4-Flash is $0.22 input / $0.66 output off-peak — about 2.5x the circulating input figure and 3.7x the circulating output figure.
Both kinds of number can be real. Resellers, routers, and promotional listings genuinely do quote lower rates, sometimes as loss-leaders, sometimes for a quantised or differently-served variant, sometimes for a limited window. The problem is comparing one of those against OpenAI's or Anthropic's list price, which is what most published comparisons do. That is not a like-for-like comparison — it is a discounted third-party rate benchmarked against a vendor rate card, and it can overstate DeepSeek's advantage by a factor of three.
Every DeepSeek, OpenAI and Anthropic figure on this page is vendor-direct list price, read off each vendor's own pricing documentation. If you are pricing a migration, insist on the same discipline — and if you are actually planning to buy through a reseller, price it against the other vendors' reseller rates too, not against their list.
How does DeepSeek's peak/off-peak pricing actually work?
Since 16 August 2026, DeepSeek bills on a two-rate schedule rather than a flat rate. The mechanics are simple but the arithmetic catches people out.
- Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday. All other hours are off-peak.
- Off-peak rates are half of peak rates across input, output, and cache reads.
That is 7 peak hours a day on 5 days a week: 35 of the 168 hours in a week, or 20.8%. Off-peak covers the other 79.2% — all weekend, plus 17 hours of every weekday.
This is where most published blends go wrong. Treating peak as 7 hours of every day (a 24-hour average) gives 29% peak and inflates every blended figure by about 7%. The Monday-to-Friday restriction is doing real work: two full weekend days at the discounted rate move the average noticeably.
| DeepSeek rate (per 1M tokens) | Off-peak (79% of week) | Peak (21% of week) | Weekly blended average |
|---|---|---|---|
| V4-Flash input (cache miss) | $0.22 | $0.44 | $0.27 |
| V4-Flash input (cache hit) | $0.007 | $0.014 | $0.0085 |
| V4-Flash output | $0.66 | $1.32 | $0.80 |
| V4-Pro input (cache miss) | $0.66 | $1.32 | $0.80 |
| V4-Pro input (cache hit) | $0.022 | $0.044 | $0.027 |
| V4-Pro output | $1.98 | $3.96 | $2.39 |
The blended column assumes traffic spread evenly across the week. If your workload is a nightly batch job or a weekend backfill, you may sit entirely off-peak and should use that column. If you serve European or US business hours, a chunk of your traffic lands in the 06:00-10:00 UTC window and your real average sits above the blend. Measure your own hour-of-day distribution before you budget — this schedule rewards moving batch work by a few hours more than almost any other optimisation on this page.
Background on how the schedule arrived: the August 2026 price change breakdown covers the transition, and the May 2026 permanent price cut covers the flat-rate era it replaced.
How much cheaper is DeepSeek V4 than GPT-5.6, vendor-direct?
All rates per 1M tokens, list price, on each vendor's own API. DeepSeek rows show off-peak / peak. OpenAI rows show short-context rates; see the long-context note below.
| Model | Input | Cached input | Output | Notes |
|---|---|---|---|---|
| DeepSeek V4-Flash | $0.22 / $0.44 | $0.007 / $0.014 | $0.66 / $1.32 | Off-peak / peak |
| DeepSeek V4-Pro | $0.66 / $1.32 | $0.022 / $0.044 | $1.98 / $3.96 | Off-peak / peak |
| GPT-5.6 Luna | $0.20 | $0.02 | $1.20 | Cost tier |
| GPT-5.6 Terra | $2.00 | $0.20 | $12.00 | Balanced tier |
| GPT-5.6 Sol | $4.00 | $0.40 | $20.00 | Flagship; promotional rate |
| GPT-6 Astra | $10.00 | $1.00 | $50.00 | New top of market, 3 Sep 2026 |
The headline multipliers, using DeepSeek's weekly blended average against GPT-5.6 Sol's live rate:
- V4-Pro vs Sol: 5.0x cheaper on input, 8.4x on output blended. Off-peak it stretches to 6.1x / 10.1x; at peak it narrows to 3.0x / 5.1x.
- V4-Pro vs GPT-6 Astra: 12.5x on input, 20.9x on output blended — 15.2x / 25.3x off-peak.
- V4-Flash vs Luna: the closest fight on the board, and DeepSeek loses half of it. Luna's $0.20 input undercuts V4-Flash at every hour of the day, including off-peak. V4-Flash wins on output off-peak ($0.66 against $1.20) and loses at peak ($1.32).
One thing the OpenAI rows hide. Every GPT model above has a second, higher rate card for long prompts: requests over 272,000 input tokens are billed at 2x input and 1.5x output for the whole request. Sol goes to $8 / $30, Astra to $20 / $75, Terra to $4 / $18, Luna to $0.40 / $1.80. DeepSeek has no equivalent threshold — its rate depends on the clock, not on prompt length. On long-context work the gap is materially wider than the table suggests.
How do the cache rates compare?
This is the most under-discussed line in every one of these rate cards, and it is where an agent workload actually lives.
OpenAI and Anthropic both price a cache read at exactly 10% of their own base input rate — that holds across Astra, Sol, Terra, Luna, Claude Opus 5, Sonnet 5, and Haiku 4.5. DeepSeek prices a cache hit at roughly 3.3% of its own input rate ($0.022 against $0.66 on V4-Pro; $0.007 against $0.22 on V4-Flash).
So DeepSeek's discount for a cache hit is about three times steeper as a ratio, on top of an input rate that is already several times lower. In absolute terms a V4-Pro cache read at $0.022 sits against $0.40 on Sol and $1.00 on Astra — roughly 18x and 45x. If you are holding a large system prompt or a long tool-call history across turns, that line does more for your bill than the headline input rate.
One honest correction to the usual framing: DeepSeek no longer has the best cache ratio on the board. Anthropic prices cache reads on Claude Fable 5.1 at 2.5% of base input ($0.25 against $10), which is a steeper proportional discount than DeepSeek's. DeepSeek still wins the absolute number by an order of magnitude — it is just no longer uniquely structured.
What does a real monthly bill look like?
Per-token rates are abstract. Take a mid-sized team running a coding agent at 50M input + 10M output tokens per month — realistic once agents run on pull requests and CI. All figures list price, vendor-direct, short context.
| Model | Monthly cost (50M in / 10M out) | vs V4-Pro blended |
|---|---|---|
| DeepSeek V4-Flash | $21.27 blended ($17.60 off-peak / $35.20 peak) | 0.33x |
| GPT-5.6 Luna | $22.00 | 0.34x |
| DeepSeek V4-Pro | $63.80 blended ($52.80 off-peak / $105.60 peak) | 1x (baseline) |
| Claude Haiku 4.5 | $100.00 | 1.6x |
| Claude Sonnet 5 | $200.00 | 3.1x |
| GPT-5.6 Terra | $220.00 | 3.4x |
| GPT-5.6 Sol | $400.00 | 6.3x |
| Claude Opus 5 | $500.00 | 7.8x |
| Claude Fable 5.1 | $1,000.00 | 15.7x |
| GPT-6 Astra | $1,000.00 | 15.7x |
Annualised, that is the difference between a ~$766/year line item on V4-Pro and ~$12,000/year on GPT-6 Astra for identical token volume. Across a team running dozens of agents, the flagship bill becomes a headcount-sized number.
Two caveats that stop this table being the whole answer. First, cheap tokens are not cheap outcomes — a model that burns more turns to reach the same result closes some of the gap on cost per solved task. Second, OpenAI makes this argument explicitly about Astra, saying it "achieves stronger results while using substantially fewer output tokens — delivering a lower estimated API cost per task than earlier models despite its higher per-token pricing." That is a claim about token efficiency, not a discount, and whether it holds depends entirely on your workload. Measure cost per completed task, not cost per million tokens, before you conclude anything from a 15x rate difference.
Is DeepSeek V4 actually good enough for coding?
Cost only matters if quality holds. The most useful evidence here is Vals AI's SWE-bench Verified leaderboard, because it runs every model on an identical minimal bash-tool-only agent harness rather than each vendor's own scaffolding — which is why its numbers differ from vendor self-reports and why they are worth more.
On that board, last updated 1 September 2026 across 86 evaluated models:
| Model | SWE-bench Verified (Vals, bash-only harness) |
|---|---|
| Claude Opus 5 | 97.00% |
| DeepSeek V4-Pro-0813 | 96.40% |
| Kimi K3 | 93.40% |
| Claude Opus 4.8 | 88.60% |
| Grok 4.5 | 86.60% |
DeepSeek's production V4-Pro-0813 checkpoint sits second of the entire field, 0.6 points behind Claude Opus 5 — while costing roughly an eighth as much per token. That is the single strongest data point for the cheap-model case, and it comes from a neutral harness rather than DeepSeek's marketing.
Read the rest of the board before you over-index on it, though. Seven of the 86 models evaluated score 95% or better, and Vals has archived this benchmark on the grounds that performance has saturated — new releases are no longer tested against it. A benchmark that cannot separate its top seven entrants is telling you the frontier has converged on this task type, not which model to ship.
The honest read: on discrete, well-scoped coding tasks — refactors, test generation, bug fixes, code review — V4-Pro is functionally interchangeable with a flagship, and you pay a fraction of the output cost. Longer autonomous runs are a different question, and one this benchmark does not answer. For the wider picture on the model family, see the DeepSeek V4 complete guide and the V4-Pro-0813 checkpoint guide.
Where does GPT-6 Astra land on cost?
Astra launched on 3 September 2026 as OpenAI's most capable model, rolling out first to enterprises in its Trusted Access Program with API and plan access following. It carries a 1,050,000-token context window, 128K max output, and a 30 April 2026 knowledge cutoff.
On price it resets the top of the market rather than the middle: $10 / $50 short context, $20 / $75 above 272K input tokens, with batch at $5 / $25 and fast mode at $20 / $100. Against DeepSeek V4-Pro's blended $0.80 / $2.39, that is 12.5x on input and 20.9x on output — and up to 25x on input if your prompts run long enough to trigger Astra's long-context tier while DeepSeek keeps billing at one flat rate.
Astra does not change the DeepSeek calculus. It widens it. The decision it forces is at the other end of the market: whether a task is valuable enough to justify twice GPT-5.6 Sol's rate. For most of the volume work this page is about — high-throughput text coding on a budget — that question does not arise. See GPT-6 Astra vs GPT-5.6 Sol for the within-OpenAI upgrade decision and GPT-6 Astra vs Claude Opus 5 for the cross-vendor one.
Where does DeepSeek V4 genuinely lose?
A cost comparison that only flatters the cheap option is not useful. Three honest caveats:
- The peak window can land on your traffic. If your users are in Europe or on US East Coast mornings, a real share of your requests fall in 06:00-10:00 UTC and you pay double. At peak, V4-Pro's advantage over Sol shrinks to 3.0x input / 5.1x output — still large, but a third of the off-peak headline. Cheap DeepSeek is partly a scheduling achievement, not purely a pricing one.
- The cheapest tier is no longer the cheapest option. Since August, GPT-5.6 Luna's $0.20 input undercuts V4-Flash at every hour. On the worked monthly bill the two are within 75 cents of each other. If V4-Flash was your default for high-volume work on price alone, that argument has expired — re-run it on your own token mix.
- Data residency. DeepSeek's API is hosted in China, which is a genuine compliance blocker for regulated industries. The counter is that V4-Pro ships open weights under an MIT licence, so self-hosting in your own region is a real option, and the API is OpenAI-compatible, so switching is a base-URL change rather than a rewrite.
The bottom line: which should you pick?
A decision rule that holds on today's verified numbers:
- High-volume, cost-is-everything: GPT-5.6 Luna or DeepSeek V4-Flash — genuinely a coin flip at $22.00 against $21.27 on the worked bill. Pick Luna if your traffic runs in DeepSeek's peak window, V4-Flash if it runs off-peak or at weekends.
- Frontier-grade coding on a budget: DeepSeek V4-Pro. Second on a neutral same-harness SWE-bench Verified board at 96.40%, at roughly an eighth of Claude Opus 5's per-token cost. This is the default pick for most engineering workloads.
- Long prompts, big repositories: DeepSeek V4-Pro, and the gap is wider than the table shows — every GPT tier doubles input pricing above 272K tokens and DeepSeek has no equivalent threshold. If you stay on OpenAI, price the long-context tier, not the headline.
- Compliance rules out China-hosted APIs: self-host V4-Pro on its MIT-licensed open weights, or move to Claude Sonnet 5 at $2 / $10 — the cheapest Western flagship-family option on the board.
- The genuinely hardest work: Claude Opus 5, Claude Fable 5.1, or GPT-6 Astra. Pay the 8-16x premium only for the narrow slice of tasks where it changes the outcome.
For most teams, the honest answer in September 2026 is unchanged by today's launch: run DeepSeek V4 as the workhorse, schedule what you can off-peak, and keep a flagship key for the slice of work that actually earns the markup. Just make sure the numbers you are comparing came from the vendors' own pricing pages — that alone will change more migration decisions than any benchmark on this page.
If the constraint is engineering capacity rather than token spend, hire vetted remote developers through Codersera to build the routing, caching, and evaluation layer that makes a multi-model setup actually cheaper in practice.
FAQ
Is DeepSeek V4 cheaper than GPT-5.6?
Yes, on vendor-direct rates. DeepSeek V4-Pro blends to about $0.80 input / $2.39 output per 1M tokens across a week, against GPT-5.6 Sol's $4 / $20 — roughly 5.0x cheaper on input and 8.4x on output, stretching to 6.1x / 10.1x off-peak. The exception is the bottom tier: GPT-5.6 Luna's $0.20 input undercuts DeepSeek V4-Flash at every hour of the day.
What are DeepSeek's peak and off-peak hours?
Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday. Every other hour, including all weekend, is off-peak. That is 35 of 168 hours a week at the peak rate, so 79% of the week is off-peak. Off-peak rates are exactly half of peak rates across input, output, and cache reads.
Why do some sites list DeepSeek V4-Flash at $0.089 per 1M tokens?
Those are third-party, reseller, or promotional listings, not DeepSeek's own rate card. DeepSeek's published price for V4-Flash is $0.22 input / $0.66 output off-peak — roughly 2.5x the circulating input figure and 3.7x the output figure. Comparing a discounted reseller rate against OpenAI's list price is not like-for-like and can overstate DeepSeek's advantage threefold.
Is GPT-5.6 Sol's promotional pricing still available?
Yes. OpenAI's pricing page states GPT-5.6 Sol's promotional pricing is available at least through 21 November 2026. The current rate is $4 input / $20 output per 1M tokens on short context, with batch at $2 / $10 and fast mode at $8 / $40. Above 272,000 input tokens the rate rises to $8 / $30 for the whole request.
How much does GPT-6 Astra cost compared to DeepSeek V4?
GPT-6 Astra costs $10 input / $50 output per 1M tokens on short context and $20 / $75 above 272K input tokens. Against DeepSeek V4-Pro's blended $0.80 / $2.39, that is about 12.5x on input and 20.9x on output — rising to roughly 25x on input for long prompts, since DeepSeek has no long-context surcharge.
Is DeepSeek V4 as good as GPT-5.6 for coding?
On discrete coding tasks, effectively yes. Vals AI's SWE-bench Verified leaderboard, which runs every model on an identical bash-tool-only harness, places DeepSeek V4-Pro-0813 second of 86 models at 96.40%, behind only Claude Opus 5 at 97.00%. Note that seven models clear 95% and Vals has archived the benchmark as saturated, so it no longer separates the top of the field.
How do DeepSeek's cache prices compare to OpenAI's and Anthropic's?
DeepSeek charges about 3.3% of its input rate for a cache hit ($0.022 against $0.66 on V4-Pro). OpenAI and Anthropic both charge exactly 10% of base input across Astra, Sol, Terra, Luna, Claude Opus 5, Sonnet 5, and Haiku 4.5. The exception is Claude Fable 5.1 at 2.5%, a steeper ratio than DeepSeek's — though DeepSeek still wins the absolute number by roughly 18-45x.
Can I self-host DeepSeek V4 to avoid the China-hosted API?
Yes. V4-Pro is released under an MIT licence with open weights, so you can run it in your own region or on-premises to satisfy data-residency requirements. The hosted API is also OpenAI-compatible, so moving between the two is a base-URL change rather than a rewrite. Self-hosting also removes the peak/off-peak schedule from your cost model entirely.