DeepSeek V4 Flash Is Retired: V4.1 Flash vs V4 Pro
Quick answer. DeepSeek V4 Flash is retired — requests to deepseek-v4-flash now serve DeepSeek V4.1-Flash, which scores 39.46 on Artificial Analysis's Intelligence Index v4.3.2 against V4 Pro's 36.00. Flash now beats Pro on the index, costs 2.5x less per task, runs twice as fast, and accepts images. Default to Flash.
This is a benchmarks-and-pricing review of DeepSeek's Flash tier, written to answer one question: should you build on it, and how does it compare to V4 Pro on the numbers that affect your bill? It is deliberately not an install walkthrough. If you want to run Flash on a Mac Studio or a GPU rig, follow our dedicated DeepSeek Flash local setup guide instead.
Important — V4 Flash is retired (September 10, 2026). Per DeepSeek's pricing page: "The legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted, but the corresponding models have been retired, their requests are served by the DeepSeek-V4.1-Flash model and billed at the Flash price."
The current model name is deepseek-flash. Nothing breaks if you keep calling the old name — but you are getting a different, better, cheaper model than the one most V4-Flash reviews describe. This review covers V4.1-Flash, the model you are actually calling.
The verdict up front, and it is the opposite of what it was in May. On the Artificial Analysis Intelligence Index v4.3.2 (a neutral third-party composite of ten evaluations, read 5 October 2026), V4.1-Flash at Max effort scores 39.46 against V4 Pro's 36.00. Flash is no longer the cheap compromise in the family — it is the higher-scoring model, at roughly 40% of Pro's cost per completed task and roughly twice the throughput. The remaining case for Pro is narrow and specific, and this review maps it.
One caveat on every number below. Artificial Analysis rebased the Intelligence Index from v4.1.1 to v4.3.2, adding Terminal-Bench 4.0 and AutomationBench-AA to the basket. Scores dropped across the entire board without any model getting worse, and there is no conversion factor between the versions. Earlier figures for this family — the "47 vs 52" pair widely quoted through mid-2026 — were v4.1.1 numbers for checkpoints that no longer exist. They cannot be compared with anything here.
Is DeepSeek Flash worth it?
For the large majority of LLM-backed product work, yes — and the case is now stronger than when Flash was the budget option. It rests on four measured facts:
- It outscores its own flagship. Intelligence Index v4.3.2 of 39.46 at Max effort, against V4 Pro 0813's 36.00.
- It is the cheaper model per completed task. Artificial Analysis measures an average $0.27 per Index task for V4.1-Flash against $0.67 for V4 Pro — roughly 2.5x.
- It is roughly twice as fast. Median output 213 tok/s versus Pro's 110 tok/s, with a 1.05s median time to first chunk against Pro's 1.67s.
- It is the multimodal one. V4.1-Flash takes image input natively (MMMU-Pro 77.0%); V4 Pro is text-only. The old "both are text-only" framing is dead.
Where Pro still wins is factual reliability, and that is now essentially the whole argument. On AA-Omniscience — which penalises confident wrong answers — Pro scores +0.83 and Flash −5.3. On Humanity's Last Exam Pro leads 41.0% to 39.2%, and on CritPt 18.0% to 14.3%. If a confident fabrication is worse for your product than no answer, that is the trade you are buying.
What is DeepSeek V4.1-Flash?
V4.1-Flash shipped on 10 September 2026 and is architecturally a bigger departure than the version number suggests. Per DeepSeek's release notes, it is a "552B-parameter MoE" with a "new Causal Encoder–Decoder architecture: just 8B active parameters for input, 16B for output." The headline specs:
- 552B total parameters, 8B active on input and 16B on output. The retired V4-Flash was 284B total / 13B active; V4 Pro is 1,600B / 49B.
- 1M-token context window, the same as Pro.
- Native multimodal input. DeepSeek's notes describe "native visual understanding" — this is where the separately-released V4-Flash-Vision-Exp went. That experimental vision model is also retired, folded into V4.1-Flash.
- MIT licence, open weights at deepseek-ai/DeepSeek-V4.1-Flash — roughly 869k downloads in the last 30 days per the Hugging Face API.
- Effort tiers. Artificial Analysis scores Max effort at 39.46 and non-reasoning at 24.67 — a 14.8-point spread, so the effort setting matters more than the model choice for most workloads.
The asymmetric active-parameter design is the reason the economics work: 8B active on the prefill means long-context input is cheap to process, while 16B on decode keeps generation quality up. That is also why Flash holds a 1M window on hardware where Pro needs a datacentre.
What does DeepSeek Flash cost?
These are the first-party rates from the official DeepSeek API pricing page, read 5 October 2026. DeepSeek now bills on a peak/off-peak split, and the old flat $0.14 / $0.28 Flash sticker is history.
| Model / tier (per 1M tokens) | Cache-hit input | Cache-miss input | Output |
|---|---|---|---|
deepseek-flash — off-peak | $0.003 | $0.15 | $0.60 |
deepseek-flash — peak | $0.006 | $0.30 | $1.20 |
deepseek-v4-pro — off-peak | $0.022 | $0.66 | $1.98 |
deepseek-v4-pro — peak | $0.044 | $1.32 | $3.96 |
Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday, excluding Chinese public holidays; everything else — including all weekends — is off-peak. Off-peak is exactly half of peak on every line. If your workload is batch and you control scheduling, shifting it out of those seven weekday hours halves the bill with no code change.
The ratio that drives model selection is stable across tiers: Pro output costs 3.3x Flash output at both peak ($3.96 vs $1.20) and off-peak ($1.98 vs $0.60). Cache-hit input is where the gap is widest — Pro's cached input is 7.3x Flash's. Any chat, agent, or RAG workload with a stable system prompt collapses input cost to near zero on Flash automatically; the API manages the cache for you.
Note also that the pricing page now lists exactly two models, deepseek-flash and deepseek-v4-pro. The legacy deepseek-chat and deepseek-reasoner aliases from the V3 era are gone from it entirely.
How does Flash score against Pro on benchmarks?
Every row below is an Artificial Analysis measurement on Intelligence Index v4.3.2, read 5 October 2026. Both models are at Max effort. We have dropped the vendor-reported SWE-bench and LiveCodeBench rows that earlier versions of this review carried: those figures were published for the April and July V4-Flash checkpoints, which no longer exist, and DeepSeek has not published equivalents for V4.1-Flash. Carrying a retired checkpoint's coding score forward would be the wrong kind of continuity.
| Metric — Intelligence Index v4.3.2 | V4.1-Flash (Max) | V4 Pro 0813 (Max) | Winner |
|---|---|---|---|
| Intelligence Index | 39.46 | 36.00 | Flash, by 3.5 |
| Average cost per Index task | $0.27 | $0.67 | Flash, 2.5x |
| Cost to run the whole Index suite | $476.89 | $1,122.27 | Flash |
| GDPval-AA (real-world work tasks) | 1600 | 1455 | Flash |
| Terminal-Bench 4.0 (agentic shell) | 26.8% | 14.1% | Flash, by 12.7pp |
| Terminal-Bench Science | 9.0% | 5.7% | Flash |
| AA-LCR (long-context reasoning) | 84.0% | 80.3% | Flash |
| SciCode | 51.9% | 51.0% | Flash (margin) |
| AA-Omniscience (hallucination-penalised) | −5.3 | +0.83 | Pro |
| Humanity's Last Exam | 39.2% | 41.0% | Pro |
| CritPt | 14.3% | 18.0% | Pro |
| MMMU-Pro (image reasoning) | 77.0% | not applicable | Flash (Pro is text-only) |
| Median output speed | 213 tok/s | 110 tok/s | Flash, 1.9x |
| Median time to first chunk | 1.05 s | 1.67 s | Flash |
Read that table as a reversal, not a refinement. Under the old V4-Flash the story was "5 points behind Pro, much cheaper." Under V4.1-Flash it is "3.5 points ahead of Pro, still much cheaper, and twice as fast." The three rows where Pro wins are all knowledge-and-reliability rows. That is a coherent, narrow moat — not a general capability lead.
The Terminal-Bench 4.0 result deserves emphasis because it inverts the old received wisdom most sharply. Agentic shell work was previously the canonical reason to escalate from Flash to Pro. On the current index component measuring exactly that, Flash scores 26.8% and Pro 14.1%. If you built a router in mid-2026 that sends long tool-call chains to Pro, that rule is now backwards.
How fast is DeepSeek Flash?
Artificial Analysis measures V4.1-Flash at a median 213 output tokens/sec against V4 Pro's 110 tok/s, with a 1.05s median time to first chunk versus Pro's 1.67s. On 100K-token prompts Flash actually measures faster still (244 tok/s), which is the asymmetric 8B-active prefill doing its job. For latency-sensitive UX — inline IDE suggestions, voice, live writing — that is the difference between "feels live" and "server is thinking." Flash is also verbose at Max effort, so cap output tokens if you are cost-sensitive.
Flash vs Pro: which should you use?
The routing rule has simplified. Default to Flash. Escalate to Pro only for factual-recall and hallucination-sensitive work — and consider whether grounded retrieval solves that better than a model swap. Reserve a closed frontier model (Claude Opus 5.5 or similar) for the hardest slice.
| Workload | Recommended | Why |
|---|---|---|
| IDE coding agent / Copilot replacement | Flash | Higher index (39.46 vs 36.00) and 3.3x cheaper output |
| Agentic shell / multi-step tool chains | Flash | Terminal-Bench 4.0 26.8% vs 14.1% — this reversed in September |
| Bulk classification / summarisation / labeling | Flash | 213 tok/s and $0.15/$0.60 off-peak |
| RAG / chatbot with stable system prompt | Flash | $0.003 cached input off-peak makes blended input near-free |
| Long-context document analysis (1M) | Flash | Same 1M window, higher AA-LCR (84.0% vs 80.3%) |
| Screenshot debugging / document OCR / vision | Flash | Only one of the two that takes images (MMMU-Pro 77.0%) |
| Latency-sensitive interactive UX | Flash | 1.9x faster output, lower TTFT |
| Privacy / on-prem / air-gapped | Flash | 16B active fits desk-class hardware; Pro needs a datacentre |
| Hallucination-sensitive enterprise lookups | Pro (+ grounding) | AA-Omniscience +0.83 vs −5.3 |
| Graduate-level recall and chained inference | Pro | HLE 41.0% vs 39.2%, CritPt 18.0% vs 14.3% |
The cost-per-task figure is the clincher and it no longer needs a caveat. Artificial Analysis spent an average $0.27 per Index task on Flash against $0.67 on Pro, and Flash scored higher. There is no capability premium being bought with the extra spend outside the three knowledge rows above. Use per-task cost rather than whole-suite totals when comparing: AA runs different task counts per model, so suite totals are not directly comparable across models — quote them only as a single model's absolute figure.
Companion guide
For the full picture on the current Flash model — architecture, deployment patterns, and benchmarks — see our DeepSeek V4.1-Flash complete guide, or the DeepSeek V4 family guide.
How do you call DeepSeek Flash from code?
The DeepSeek API is OpenAI-compatible: endpoint https://api.deepseek.com/v1/chat/completions, model name deepseek-flash, Bearer auth. Prompt caching is automatic when your prompt prefix repeats; the usage object reports how much was billed at the cached rate. For a full local deployment, the setup guide covers MLX/vLLM/llama.cpp.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["DEEPSEEK_API_KEY"],
base_url="https://api.deepseek.com/v1",
)
SYSTEM_PROMPT = """You are a senior code reviewer. Review the diff for
correctness, security, and performance issues. Return JSON:
{ issues: [{severity, file, line, message}] }"""
def review(diff: str) -> str:
r = client.chat.completions.create(
model="deepseek-flash", # "deepseek-v4-flash" still routes here
messages=[
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": diff},
],
response_format={"type": "json_object"},
temperature=0.2,
)
u = r.usage
hit = getattr(u, "prompt_cache_hit_tokens", 0)
miss = getattr(u, "prompt_cache_miss_tokens", 0)
print(f"cache hit {hit} / miss {miss} / out {u.completion_tokens}")
return r.choices[0].message.contentThe trick to maximise cache hits: keep the system prompt byte-stable and put long stable context (style guide, docs, schema) before short variable context (the diff, the user message). DeepSeek's caching is prefix-based; any change to the front of the prompt invalidates the cached prefix. And if your job is schedulable, run it outside 01:00–04:00 and 06:00–10:00 UTC on weekdays to halve every line of the bill.
What should you migrate, and when?
Three concrete actions, in order:
- Rename the model string to
deepseek-flash. The legacy name works, but pinning to a retired alias is how you end up surprised by a silent model swap. One-line change. - Re-test any Flash-to-Pro escalation rule you wrote before September. If it escalates long tool-call chains or agentic shell work to Pro, it is now routing your hardest traffic to the weaker and more expensive model.
- Re-price your forecast. Flash is no longer $0.14/$0.28 flat. Peak output is $1.20/M — more than 4x the old sticker — and a workload that assumed the flat rate will overrun. The off-peak window is the mitigation.
What has not changed: V4 Pro is still on the API. DeepSeek's changelog states that "in response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged." An earlier note had signalled that Pro requests would be routed to V4.1-Flash; that was reversed. You do not need to migrate off Pro — you just no longer have a strong reason to be on it.
FAQ
Is DeepSeek V4 Flash still available?
No. DeepSeek's pricing page states that V4-Flash and V4-Flash-Vision-Exp are retired and that requests to those legacy model names are served by DeepSeek-V4.1-Flash and billed at the Flash price. The current model name is deepseek-flash. Your code keeps working; the model behind it changed on 10 September 2026.
Is DeepSeek Flash worth it over V4 Pro?
Yes, for most workloads — and now on capability as well as cost. V4.1-Flash scores 39.46 on Artificial Analysis's Intelligence Index v4.3.2 against V4 Pro's 36.00, at $0.27 per Index task versus $0.67, at roughly twice the throughput. Escalate to Pro only for factual recall and hallucination-sensitive work, where Pro leads on AA-Omniscience (+0.83 vs −5.3).
How much does DeepSeek Flash cost per million tokens?
Off-peak: $0.15 cache-miss input, $0.003 cache-hit input, $0.60 output. Peak: $0.30, $0.006 and $1.20. Peak is 01:00–04:00 and 06:00–10:00 UTC on weekdays excluding Chinese public holidays; off-peak is exactly half of peak on every line. The old flat $0.14 / $0.28 V4-Flash rate no longer applies.
What does DeepSeek Flash score on the AA Intelligence Index?
39.46 at Max effort on Intelligence Index v4.3.2, and 24.67 in non-reasoning mode — a 14.8-point spread, so always state the effort tier with the score. V4 Pro 0813 scores 36.00 at Max effort. Any "47" you see quoted is a v4.1.1 figure for the retired V4-Flash checkpoint and is not comparable.
Does DeepSeek Flash support images?
Yes. V4.1-Flash has native multimodal input per DeepSeek's release notes, and Artificial Analysis measures it at 77.0% on MMMU-Pro. This is new — the retired V4-Flash was text-only, and the separate V4-Flash-Vision-Exp experiment was folded into V4.1-Flash. V4 Pro remains text-only, so Flash is the multimodal option in the family.
When should I use Pro instead of Flash?
Three cases, all knowledge-shaped: hallucination-sensitive factual lookups (AA-Omniscience +0.83 vs −5.3), graduate-level recall and chained inference (Humanity's Last Exam 41.0% vs 39.2%), and hard physics-style reasoning (CritPt 18.0% vs 14.3%). Agentic tool chains are no longer on this list — Flash now leads Terminal-Bench 4.0 by 12.7 points.
How big is V4.1-Flash compared to the old V4-Flash?
Larger in total and smaller in active parameters. V4.1-Flash is a 552B-parameter MoE using a causal encoder–decoder design with 8B active on input and 16B on output; the retired V4-Flash was 284B total with 13B active. Both carry a 1M-token context window. The asymmetric design is why long-context prefill is cheap on V4.1-Flash.
Where do I learn how to run DeepSeek Flash locally?
This review intentionally does not cover installation. Follow our dedicated DeepSeek Flash local setup guide for hardware requirements, quantisation choices, and MLX/vLLM/llama.cpp serving. The weights are MIT-licensed on Hugging Face.
What is the verdict for engineering teams?
The default for LLM-backed product work has not changed, but the reason has. DeepSeek Flash is the model you build on — not because it is an acceptable compromise, but because on Intelligence Index v4.3.2 it is the better of DeepSeek's two models at 39.46 against Pro's 36.00, while costing 2.5x less per completed task, running ~1.9x faster, and being the only one that takes images. Pro has narrowed to a knowledge-reliability specialist.
The harder part is execution: most of the savings live in caching, off-peak scheduling, routing, and batching — engineering work that never shows up in a benchmark. And as the September reversal shows, a routing rule written four months ago can quietly become wrong. If you're hiring vetted remote developers experienced with DeepSeek, LLM cost engineering, and production AI pipelines, codersera.com/hire matches you with senior, remote-ready talent in TypeScript, Python, Go, Rust, and Node who have shipped this kind of work before.
Sources and further reading
- DeepSeek API pricing (vendor, primary — peak/off-peak rates and the V4-Flash retirement note)
- DeepSeek-V4.1-Flash release notes, 10 September 2026 (vendor, primary)
- DeepSeek API changelog (vendor, primary — V4 Pro continuation)
- DeepSeek-V4.1-Flash on Hugging Face (vendor, MIT weights)
- Artificial Analysis: DeepSeek V4.1-Flash (neutral, third-party)
- Artificial Analysis: DeepSeek V4 Pro (neutral, third-party)
- Codersera: DeepSeek V4.1-Flash complete guide
- Codersera: Run DeepSeek Flash locally — full setup guide
- Codersera: DeepSeek V4 complete guide (2026) (pillar)
- Codersera: Best open-source LLM in 2026 (where Flash sits against the wider open-weights field)