DeepSeek V4 Pro vs Flash: Which to Use in 2026
On April 24, 2026, DeepSeek shipped two models on the same day: DeepSeek V4-Pro (1.6T total / 49B active parameters) and DeepSeek V4-Flash (284B total / 13B active). Both shared an architecture family, both were trained on 32T+ tokens, and both shipped a legitimate 1M token context window. Five months later the comparison has changed shape twice: V4-Flash is retired, replaced by the larger and better V4.1-Flash, and the Flash tier now outscores Pro on the neutral index rather than trailing it. Flash output still costs 3.3x less than Pro's; it is no longer buying you less capability to get that.
This is the DeepSeek V4 Pro vs Flash comparison engineering leaders actually need: not a press-release recap, but a hard look at where the 3.5-point Artificial Analysis Intelligence Index gap — now in Flash's favour — shows up in production, where it doesn't, and how to map each variant to real workloads. We'll cover benchmarks, cost-per-task economics, provider speed tiers, local deployment feasibility, and a use-case decision tree. If you're also weighing DeepSeek V4 against Western frontier models, see our companion pieces on DeepSeek V4 vs Claude Opus 4.7 and DeepSeek V4 vs GPT-5.5 Pro.
Short version: Flash is the default for most production code, RAG, tool-calling, long-context and vision workloads in 2026. Pro is the right choice for hallucination-sensitive factual recall and graduate-level knowledge work. Agentic tool chains used to be a reason to pick Pro; on the current index they are a reason to pick Flash. Read on for the data behind that claim.
Want the full picture? Read our continuously-updated DeepSeek V4 complete guide — benchmarks, pricing, deployment patterns, and how it compares to GPT-5.5 and Claude Opus 4.7.
Update — V4-Flash is retired (September 10, 2026). Per DeepSeek's pricing page, the current Flash model name is deepseek-flash, and "the legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted, but the corresponding models have been retired, their requests are served by the DeepSeek-V4.1-Flash model and billed at the Flash price."
So the live comparison is V4-Pro vs V4.1-Flash, and that is what this page now measures. V4.1-Flash is a 552B-parameter MoE with 8B active on input and 16B on output, a 1M context, and native image input — see our V4.1-Flash complete guide.
Update — V4-Pro reached GA on August 13, 2026, and is still on the API. The April 24 launch was a preview. The production checkpoint is V4-Pro-0813, which added reasoning-effort levels (low / high / max), a native OpenAI Responses API with one-click Codex setup, and DSpark speculative decoding. Peak/off-peak pricing took effect on August 16 — see the August 2026 price change breakdown. A September note had signalled that V4-Pro requests would be routed to V4.1-Flash; that was reversed. DeepSeek's changelog states it "decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged."
The two variants in 60 seconds
Both models share architecture: a hybrid attention stack combining Compressed Sparse Attention (CSA), DeepSeek Sparse Attention (DSA), and Heavily Compressed Attention (HCA), alongside Manifold-Constrained Hyper-Connections (mHC) and the Muon optimizer. Critically, Flash was never a distillation of Pro — it is a separate training run within the same family, post-trained with SFT, GRPO, and on-policy distillation. V4.1-Flash goes further and changes architecture outright, to a causal encoder–decoder with asymmetric active parameters (8B on input, 16B on output), which is what lets it process a 1M-token prefill cheaply. It also added native image input; V4-Pro remains text-only, so the modality advantage in this family now belongs to Flash.
| Spec | V4-Pro 0813 | V4.1-Flash (current) | V4-Flash 0731 (retired) |
|---|---|---|---|
| Total / active params | 1.6T / 49B | 552B / 8B in, 16B out | 284B / 13B |
| Context window | 1M tokens | 1M tokens | 1M tokens |
| Architecture | MoE, CSA + DSA + HCA | Causal encoder–decoder MoE | MoE, CSA + DSA + HCA |
| Modality | Text-only | Text + image | Text-only |
| Reasoning modes | low / high / max | non-reasoning / max | high / xHigh |
| Licence | MIT | MIT | MIT |
| Released | 13 Aug 2026 (GA) | 10 Sep 2026 | 31 Jul 2026 |
| FLOPs at 1M ctx vs V3.2 | 27% | — (not published) | 10% |
| KV cache at 1M ctx vs V3.2 | 10% | — (not published) | 7% |
The compression rows are the quiet story: Flash was never just smaller, it was more aggressively compressed at long context. At 1M tokens the retired V4-Flash used roughly a third of Pro's FLOPs and 70% of Pro's KV cache, which is why Flash can hold a 1M-token conversation on a single Mac Studio while Pro requires an 8× H100 server. V4.1-Flash pushes the same idea into the architecture itself — 8B active parameters on the input pass — though DeepSeek has not published comparable FLOPs and KV-cache ratios for it.
V4 Pro vs Flash pricing
Pricing is where the architectural choice cashes out, and it has moved twice since launch. The 75%-off V4-Pro promotion that ran through mid-2026 is history; since 16 August 2026 both models bill on a peak/off-peak split, and the V4.1-Flash release on 10 September cut the Flash tier again. These are the current first-party rates from the official DeepSeek pricing page, read 5 October 2026.
| Model / tier (per 1M tokens) | Cache-hit input | Cache-miss input | Output |
|---|---|---|---|
deepseek-flash — off-peak | $0.003 | $0.15 | $0.60 |
deepseek-flash — peak | $0.006 | $0.30 | $1.20 |
deepseek-v4-pro — off-peak | $0.022 | $0.66 | $1.98 |
deepseek-v4-pro — peak | $0.044 | $1.32 | $3.96 |
Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday, excluding Chinese public holidays; all other hours — including every weekend — are off-peak, and off-peak is exactly half of peak on every line. If your workload is batch and you control scheduling, moving it out of those seven weekday hours halves the bill with no code change.
The ratio that drives model selection is stable across tiers: Pro output costs 3.3× Flash output at both peak ($3.96 vs $1.20) and off-peak ($1.98 vs $0.60). Cache-hit input is where the gap is widest, at 7.3×. Any chat, agent or RAG workload with a byte-stable system prompt collapses input cost to near zero on Flash automatically; the API manages the cache for you.
One warning for anyone whose forecast predates August: the old flat $0.14 / $0.28 Flash sticker no longer exists. Peak Flash output is now $1.20/M, more than 4× that rate, and a budget built on the flat price will overrun. The off-peak window is the mitigation.
| Model / tier (from 2026-08-16) | Cache hit | Cache miss (in) | Output |
|---|---|---|---|
| V4-Flash — off-peak | $0.007 | $0.22 | $0.66 |
| V4-Flash — peak | $0.014 | $0.44 | $1.32 |
| V4-Pro — off-peak | $0.022 | $0.66 | $1.98 |
| V4-Pro — peak | $0.044 | $1.32 | $3.96 |
Blended, that is roughly 1.96x on input, 2.94x on output and 7.84x on cache reads. The Pro-vs-Flash ratio survives the change — Pro output stays 3x Flash output at both tiers — so the decision below does not flip. Your bill does.
Cost-per-task: the $0.27 vs $0.67 number
Sticker pricing only tells you what a token costs. The number that matters is what a completed task costs. Artificial Analysis publishes an average cost per Intelligence Index task, which is the right unit here — AA runs different task counts per model, so whole-suite dollar totals are not comparable across models and should only be quoted as a single model's absolute figure.
These figures are Intelligence Index v4.3.2, read 5 October 2026. Note that AA rebased the index from v4.1.1, adding Terminal-Bench 4.0 and AutomationBench-AA to the basket; scores fell across the whole board without any model getting worse, and there is no conversion factor. The "$113 vs $1,071" and "47 vs 52" pairs that circulated through mid-2026 were v4.1.1 figures for checkpoints that no longer exist.
| Model (Max effort) | Cost per Index task | Whole-suite cost | Index v4.3.2 |
|---|---|---|---|
| DeepSeek V4.1-Flash | $0.27 | $476.89 | 39.46 |
| DeepSeek V4-Pro 0813 | $0.67 | $1,122.27 | 36.00 |
| DeepSeek V4-Flash 0731 (retired) | $0.22 | $474.19 | 34.33 |
| Claude Opus 5.5 (Max) — closed, for reference | $5.98 | $8,708.20 | 57.62 |
Read that as a reversal, not a discount. Flash costs 2.5× less per completed task than Pro and scores 3.5 index points higher. There is no capability premium being bought with the extra Pro spend outside the knowledge benchmarks covered below. The previous framing — "pay 9.5× more for 5 more index points" — described a trade-off that no longer exists in either direction.
The cross-lab comparison still favours DeepSeek heavily. Claude Opus 5.5 at Max effort scores 57.62, an 18.2-point lead over V4.1-Flash, but costs 22× more per Index task ($5.98 vs $0.27). That gap is real and it is concentrated in agentic work; whether it is worth 22× depends entirely on how much a failed task costs you.
One methodological caution that still applies: cheaper tokens do not always mean a cheaper task. On LiveBench's agentic-coding split, the retired V4-Flash cost $0.0596 per successful task against V4-Pro's $0.0442 — roughly 35% more per completed task despite billing far less per output token, because it thought longer and retried more to land the same result. Those specific figures describe retired checkpoints and should not be read forward to V4.1-Flash, which has not been re-measured on that split. The lesson transfers even though the numbers do not: per-token sticker is the wrong unit for agentic work. Measure cost-per-success on your own traffic.
DeepSeek Flash benchmarks vs Pro: where the gap is and isn't
Every row below is an Artificial Analysis measurement on Intelligence Index v4.3.2, read 5 October 2026, with both models at Max effort. Stating the effort tier matters: V4.1-Flash scores 39.46 at Max and 24.67 in non-reasoning mode, a 14.8-point spread wider than the gap between the two models.
We have dropped the vendor-reported rows that earlier versions of this page carried — SWE-bench Verified, LiveCodeBench, Codeforces, MMLU-Pro, HMMT, IMOAnswerBench, SimpleQA-Verified, MRCR and BrowseComp. Those were published for the April and July V4-Flash checkpoints, which are retired, and DeepSeek has not published equivalents for V4.1-Flash. Terminal-Bench 2.0 is likewise gone: v4.3.2 scores Terminal-Bench 4.0 instead, and the two are not comparable.
| Metric — Index v4.3.2 | V4.1-Flash (Max) | V4-Pro 0813 (Max) | Winner |
|---|---|---|---|
| Intelligence Index | 39.46 | 36.00 | Flash, by 3.5 |
| Cost per Index task | $0.27 | $0.67 | Flash, 2.5× |
| GDPval-AA (real-world work) | 1600 | 1455 | Flash |
| Terminal-Bench 4.0 (agentic shell) | 26.8% | 14.1% | Flash, by 12.7pp |
| Terminal-Bench Science | 9.0% | 5.7% | Flash |
| AA-LCR (long-context reasoning) | 84.0% | 80.3% | Flash |
| SciCode | 51.9% | 51.0% | Flash (margin) |
| MMMU-Pro (image reasoning) | 77.0% | n/a (text-only) | Flash |
| AA-Omniscience (hallucination-penalised) | −5.3 | +0.83 | Pro |
| AA-Omniscience accuracy | 46.4% | 49.1% | Pro |
| Humanity's Last Exam | 39.2% | 41.0% | Pro |
| CritPt | 14.3% | 18.0% | Pro |
| Median output speed | 213 tok/s | 110 tok/s | Flash, 1.9× |
| Median time to first chunk | 1.05 s | 1.67 s | Flash |
Coding and engineering work: Flash leads
On the index components that proxy engineering work, Flash is ahead rather than behind. GDPval-AA, which scores real-world occupational tasks, puts Flash at 1600 against Pro's 1455. SciCode is effectively tied (51.9% vs 51.0%). DeepSeek has not published SWE-bench or LiveCodeBench figures for V4.1-Flash, so we no longer quote them — the widely-circulated 79.0 / 80.6 pair described the retired checkpoint. For your TypeScript, Python, or Go agentic coding loops, Flash is now the rational default on capability as well as cost.
Hard reasoning: Pro's narrow win
This is where Pro's 49B active parameters still earn their keep, and the margins are small but consistent. Humanity's Last Exam: 41.0% for Pro against 39.2% for Flash. CritPt, which tests research-level physics reasoning: 18.0% against 14.3%. Both reward graduate-level recall and long chained inference over a large parameter base. If your workload is frontier-difficulty reasoning rather than engineering throughput, Pro is defensible — at 2.5× the cost per task for roughly two points.
Long context: the edge flipped to Flash
This reversed with V4.1-Flash. On AA-LCR, Artificial Analysis's long-context reasoning evaluation, Flash scores 84.0% and Pro 80.3%. Flash is also measurably faster at length — AA records 244 median output tok/s on 100K-token prompts, above its own 213 tok/s median, which is the 8B-active input pass doing its job. For RAG, contract review and codebase navigation at 1M tokens, Flash is both the better and the cheaper call. The old MRCR comparison (83.5 vs 78.7) described the retired checkpoint and is not part of the current index.
Agentic and tool use: this is the reversal
Agentic shell work used to be the canonical reason to escalate from Flash to Pro. It is now the strongest single reason not to. On Terminal-Bench 4.0 — the agentic benchmark v4.3.2 actually scores — Flash gets 26.8% and Pro 14.1%, a 12.7-point gap in Flash's favour. Terminal-Bench Science follows the same direction (9.0% vs 5.7%), and so does GDPval-AA (1600 vs 1455).
If you wrote a router before September 2026 that sends long tool-call chains to Pro, that rule is now backwards — it is routing your hardest traffic to the weaker and more expensive model. The old HN read on Flash from gertlabs still rings true as a description of its style, though: "not a smart model on the first try, but it makes up for it over the course of a session." v4.3.2 weights exactly that kind of multi-step recovery more heavily than v4.1.1 did, which is part of why the ordering changed.
Knowledge and hallucination: Pro's clear win
This is the one place the old conclusion survives intact. On AA-Omniscience, which penalises confident wrong answers rather than merely counting correct ones, Pro scores +0.83 and Flash −5.3 — the difference between a model that roughly breaks even on the penalty and one that does not. Raw accuracy is 49.1% for Pro against 46.4% for Flash, and both hallucinate at high rates on the questions they get wrong (94.8% and 96.5% respectively).
For factual customer-facing systems, regulated industries, or any workflow where a confident wrong answer is worse than no answer, this is the deciding metric — and it is worth asking whether grounded retrieval fixes it more cheaply than a 2.5× model upgrade. Neither model is a substitute for retrieval against a source you control.
Speed and providers
Headline numbers from Artificial Analysis at Max effort, read 5 October 2026: V4.1-Flash at a median 213 output tokens/sec, V4-Pro at 110 t/s, with time to first chunk of 1.05s and 1.67s respectively. Flash is roughly 1.9× faster. The provider tables below were measured on the earlier checkpoints and are kept because the spread between providers is the durable lesson — treat the absolute numbers as historical, not current.
| Provider | V4-Pro Output t/s | V4-Pro TTFT (s) |
|---|---|---|
| Fireworks | 169.3 | 28.06 |
| Together.ai | 48.3 | 0.99 |
| Novita | 36.0 | 123.41 |
| SiliconFlow | 35.8 | 124.21 |
| DeepSeek 1st party | 35.6 | 1.82 |
| DeepInfra (FP4) | 32.3 | 1.27 |
| Provider | V4-Flash Output t/s | V4-Flash TTFT (s) |
|---|---|---|
| Novita | 85.5 | 67.23 |
| SiliconFlow (FP8) | 83.7 | 68.42 |
| DeepSeek | 81.3 | 70.07 |
The durable point is that provider choice moves throughput by 5× and TTFT by 100× — a far larger effect than the Flash-vs-Pro difference. Benchmark your actual provider rather than trusting a model-level number. On the first-party API as measured today, Flash wins both axes outright, so the Fireworks-Pro workaround that made sense in May is no longer needed.
Local deployment
V4-Pro: datacenter-only
Pro requires 8× H100 or H200 server class hardware. At Q4 quantization the weights still occupy roughly 800GB. This is not homelab-feasible. Notably, DeepSeek positioned Pro as the first DeepSeek model optimized for Huawei Ascend 950 silicon, which is the geopolitical subtext of the launch.
V4-Flash: real local options
Flash has a hard floor of 90GB pooled memory. Below that, you'll page from disk and the model will hallucinate from incomplete context. Above that floor:
| Hardware | Approx. cost | Throughput (Q4 / Q4_K_M) |
|---|---|---|
| Mac Studio M4 Max 192GB | $5,999 | 25-35 t/s (MLX) |
| RTX PRO 6000 96GB | ~$8,500 | 45-60 t/s |
| Dual H100 80GB | ~$50,000 | 60-90 t/s |
Day-zero tooling: vLLM (FP4/FP8), SGLang, MLX (community Flash port), and llama.cpp (antirez fork). Community quants live at unsloth/DeepSeek-V4-Flash and tecaprovn/deepseek-v4-flash-gguf. AWQ/INT4 community ports are still in flight at the time of writing.
When to use V4 Flash
The default answer for production workloads in 2026 is Flash. Specifically:
- Bulk classification, summarization, and data labeling. The 3.3× output-cost difference is decisive at millions of completions per day, and Flash now scores higher than Pro on the index overall.
- Code autocomplete and mid-tier coding agents. Cursor- or Copilot-style replacements: gertlabs on HN argued Flash is the right pick here, citing speed and the fact that it self-corrects across a session.
- Customer service chatbots and RAG over enterprise corpus. AA-LCR of 84.0% at a 1M window beats Pro's 80.3%, and cache-hit input at $0.003 / 1M off-peak makes high-traffic chat economically viable.
- Agentic tool-calling pipelines, including deep chains. Flash leads Terminal-Bench 4.0 by 12.7 points (26.8% vs 14.1%). This is the bullet that moved from Pro's column to Flash's in September.
- Long-context analysis. AA-LCR 84.0% against Pro's 80.3%, same 1M window, and faster at 100K-token prompts.
- Image input — screenshot debugging, document OCR, design-to-code. Flash is the only model in the family that takes images (MMMU-Pro 77.0%).
- Speed-prioritized interactive UX. 213 tok/s against Pro's 110, and 1.05s to first chunk against 1.67s.
- Privacy-sensitive workloads. Mac Studio runnable means you can keep the workload entirely on-device for legal, healthcare, or finance use cases.
When to use V4 Pro
Pro is the right call when the cost of being wrong outweighs the cost of being expensive. Specifically:
- Factual recall and world knowledge. AA-Omniscience +0.83 against Flash's −5.3 is the single clearest delta left. If factual correctness is the product, Pro is the safer base — paired with grounded retrieval.
- Hallucination-sensitive enterprise workflows. Regulated industries, legal review, medical research, financial analysis: a confident fabrication costs more than the 2.5× per-task premium.
- Research-grade reasoning. Humanity's Last Exam 41.0% vs 39.2% and CritPt 18.0% vs 14.3%. Narrow margins, but they point the same way.
- Text-only by policy. If your compliance posture forbids a vision-capable endpoint, Pro is the text-only option in the family.
Real-world reception
The community read on V4 has been remarkably consistent across independent voices. Simon Willison summed it up: "DeepSeek V4 - almost on the frontier, a fraction of the price." Latent Space's digest emphasized two things: that "Flash@max ≈ Pro@high on reasoning tasks" and that V4 represents "legit 1M context for pennies."
"DeepSeek V4 Flash is the model to pay attention to here. It's cheap, effective, and REALLY fast." - HN user gertlabs
The same commenter was sharper on Pro: "The Pro model is slow, not much better in coding reasoning so far when it works, and honestly too unreliable and rate limited to be of much use, currently." That is one user's experience on the DeepSeek 1st-party endpoint, where Pro caps at 35 t/s and is heavily rate-limited; on Fireworks, the speed complaint goes away. The reliability complaint is a function of Pro's deeper agentic chain failures, which the benchmarks confirm.
On Flash's agentic behavior, gertlabs added a useful nuance: "Not a smart model on the first try, but it makes up for it over the course of a session." That matches the benchmark profile: Flash is more error-prone per step, but its self-correction in chat contexts is strong.
Decision tree
If you only have 30 seconds, here is the decision flow as prose:
Start: Is your workload factual recall or regulated/hallucination-sensitive? → Pro.
Otherwise, does your agent chain 8+ tool calls per turn or do web-search-heavy research? → Pro.
Otherwise, is your workload code autocomplete, RAG, classification, summarization, single-step tool calls, or chat? → Flash.
Edge case: Need frontier reasoning at low latency? → Pro on Fireworks (169 t/s, 28s TTFT) for batch; Pro on Together.ai for interactive (sub-1s TTFT).
Edge case: Need local / on-device? → Flash on Mac Studio M4 Max 192GB minimum. Pro is datacenter-only.
What this means for engineering teams
The economic shift is no longer a trade-off, which is the headline. On Intelligence Index v4.3.2, V4.1-Flash completes an average Index task for $0.27 against Pro's $0.67 — and scores 3.5 points higher doing it. Against the closed frontier, Flash is 22× cheaper per task than Claude Opus 5.5 ($0.27 vs $5.98) for 18.2 fewer index points. That is a genuine capability gap, concentrated in agentic work, and it is the honest reason some workloads still belong on a frontier model. It is not a reason to pay 2.5× inside DeepSeek's own family for a lower score.
For engineering leaders, this means three concrete things. First, your default model for new internal tooling should be Flash, with Pro reserved for the knowledge-reliability workloads above. Second, audit any Flash-to-Pro escalation rule written before September 2026 — if it escalates agentic or long-context work, it is now sending your hardest traffic to the weaker model. Third, re-price your forecast: the flat $0.14 / $0.28 Flash rate is gone, peak output is $1.20/M, and scheduling batch work outside 01:00–04:00 and 06:00–10:00 UTC on weekdays halves every line.
If you're staffing a team to actually integrate these models, Codersera connects companies with vetted senior engineers who have shipped LLM-backed systems in production. Whether you need a Python developer for ML pipelines, a TypeScript developer for agentic frontends, a Go developer for high-throughput inference proxies, a Node.js developer for orchestration layers, or a Rust developer for performance-critical model serving, you can scope and start within days. Learn more about our services, why teams choose us, or browse our AI engineering blog for more deep dives.
FAQ
Is DeepSeek Flash distilled from V4 Pro?
No. The retired V4-Flash was an independent training run within Pro's architecture family — same hybrid attention design (CSA + DSA + HCA), same Muon optimizer, same 32T+ token mix, but not a distillation. The current V4.1-Flash diverges further still: DeepSeek describes it as a "new Causal Encoder–Decoder architecture" with 8B active parameters on input and 16B on output, which is structurally unlike Pro rather than a shrunken copy of it.
How much does DeepSeek Flash actually cost in production?
Off-peak: $0.15 cache-miss input, $0.003 cache-hit input, $0.60 output per 1M tokens. Peak: $0.30, $0.006 and $1.20. Peak is 01:00–04:00 and 06:00–10:00 UTC on weekdays excluding Chinese public holidays, and off-peak is exactly half of peak on every line. The old flat $0.14 / $0.28 rate was V4-Flash's and no longer applies. Artificial Analysis measures an average $0.27 per Index task.
What hardware do I need to run V4 Flash locally?
The hard floor is 90GB of pooled memory. Practical configurations: Mac Studio M4 Max 192GB ($5,999, 25-35 t/s on MLX Q4_K_M), RTX PRO 6000 96GB (~$8,500, 45-60 t/s), or dual H100 80GB (~$50K, 60-90 t/s). Below 90GB you'll page from disk and accuracy collapses.
Can I run V4 Pro locally?
Realistically, no. Pro requires 8× H100 or H200 class hardware. At Q4 quantization the weights still occupy roughly 800GB. This is a datacenter workload.
How does DeepSeek compare to the closed frontier now?
On Intelligence Index v4.3.2, V4.1-Flash scores 39.46 and V4-Pro 36.00 against Claude Opus 5.5 (Max) at 57.62 and GPT-5.5 (Xhigh) at 38.36 — so Flash edges GPT-5.5 while costing roughly a tenth as much per Index task, and both DeepSeek models sit well behind Anthropic's current flagship. Full breakdowns: DeepSeek V4 vs Claude Opus 4.7 and DeepSeek V4 vs GPT-5.5 Pro.
Which provider should I use?
Benchmark your own. Provider choice historically moved V4-Pro throughput by about 5× and time-to-first-token by two orders of magnitude — a far bigger effect than the model choice. On Artificial Analysis's current first-party measurements Flash leads on both throughput (213 vs 110 tok/s) and latency (1.05s vs 1.67s), so the third-party workarounds that made sense mid-year are no longer necessary.
Does DeepSeek Flash hallucinate more than V4 Pro?
Yes, and it is the clearest remaining reason to choose Pro. On AA-Omniscience, which penalises confident wrong answers, Pro scores +0.83 against Flash's −5.3, with raw accuracy of 49.1% versus 46.4%. Both hallucinate at high rates on the questions they get wrong (94.8% and 96.5%), so for hallucination-sensitive workloads Pro is the safer base — but neither is a substitute for grounded retrieval.
What happened to the V4 Pro 75% discount?
It is gone, superseded by the peak/off-peak structure that took effect at 16:00 UTC on 16 August 2026. V4-Pro now bills $0.66 input / $1.98 output per 1M off-peak and $1.32 / $3.96 at peak; the $0.435 / $0.87 standing rate from May 2026 and the earlier $1.74 / $3.48 list rates are both historical. The Flash-vs-Pro output ratio is 3.3× at either tier.
When should I just use DeepSeek Flash for everything?
Unless factual reliability is the product. Flash now scores higher than Pro on Intelligence Index v4.3.2 (39.46 vs 36.00) while costing 2.5× less per task, so code, RAG, classification, summarisation, agentic tool chains, long-context analysis, vision and chat all have a positive case for Flash rather than a cost-driven compromise. Keep Pro for hallucination-sensitive factual lookups and frontier-difficulty reasoning.
Does DeepSeek V4 support images or audio?
Images, yes — on Flash only. V4.1-Flash has native multimodal input per DeepSeek's release notes and scores 77.0% on MMMU-Pro. This is new: the retired V4-Flash was text-only, and the separate V4-Flash-Vision-Exp experiment was folded into V4.1-Flash. V4-Pro remains text-only. Neither model accepts audio.
Sources and further reading
- DeepSeek V4-Pro on Hugging Face
- DeepSeek V4.1-Flash on Hugging Face (MIT weights, current Flash model)
- DeepSeek V4.1-Flash release notes (10 September 2026)
- DeepSeek API changelog (V4-Flash retirement, V4-Pro continuation)
- Artificial Analysis — V4.1-Flash
- Codersera: DeepSeek V4.1-Flash complete guide
- DeepSeek V4 collection
- DeepSeek API pricing
- DeepSeek V4 release notes
- Artificial Analysis - V4-Pro
- Artificial Analysis - V4-Flash
- AA - V4-Pro providers
- AA - V4-Flash providers
- AA - DeepSeek is back among the leading open-weights models
- OpenRouter V4-Flash
- OpenRouter V4-Pro
- Hacker News - V4 main thread
- Hacker News - V4 technical paper
- Hacker News - Day 0 SGLang
- Latent Space digest
- Simon Willison on DeepSeek V4
- compute-market local hardware guide
- InsiderLLM V4 Pro vs Flash guide
- BuildFastWithAI - DeepSeek V4 Flash review
- OfficeChai V4 benchmarks and pricing
- DataCamp DeepSeek V4 article
- BenchLM V4-Flash
- NVIDIA - Build with DeepSeek V4 on Blackwell
For more LLM and AI engineering deep dives, see the Codersera blog, the AI tag, or our FAQs on engaging engineering talent.