DeepSeek V4-Pro 0813 Benchmarks: SWE-bench, LiveBench
DeepSeek published the 0813 checkpoint to Hugging Face at 16:28 UTC on 13 August 2026, ending the preview that began on 24 April. Ten days on, enough independent measurement exists to say something firm — and the picture splits sharply depending on which benchmark you look at.
This page collects every DeepSeek V4-Pro 0813 benchmark result we could verify against a primary source, names who measured each one, and flags where a vendor claim and an independent run disagree. All figures re-checked 23 August 2026.
What is DeepSeek V4-Pro 0813?
0813 is the general-availability checkpoint of DeepSeek V4-Pro. It is not a new base model — it is the production release of the architecture described in arXiv 2606.19348, with the DSpark speculative-decoding module folded into the shipped weights.
| Property | DeepSeek V4-Pro 0813 |
|---|---|
| Released | 13 August 2026, 16:28 UTC |
| Parameters | 1.6T total / 49B activated (MoE, 6 experts per token) |
| Precision | FP8, with MoE experts in FP4 (expert_dtype: fp4) |
| Context | 1,048,576 tokens |
| Max output | 384K recommended at high / max effort |
| Weights | 892.8 GB across 92 safetensors files |
| Licence | MIT |
| Modality | Text only |
Sizes and dates above are read straight from the Hugging Face API for deepseek-ai/DeepSeek-V4-Pro-0813. The pre-DSpark DeepSeek-V4-Pro repo is 864.7 GB across 91 files, so the extra ~28 GB in 0813 is the draft module, not a larger base network.
One number worth sitting with: on 23 August 2026 the Pro-0813 repo shows 54,566 downloads, while V4-Flash-0731 shows 2,976,281. The open-weights community is overwhelmingly on Flash. Pro is an API model that happens to have downloadable weights.
What does DeepSeek V4-Pro score on SWE-bench Verified?
96.40%. That places it second of 82 models on Vals AI's SWE-bench Verified board (updated 19 August 2026), and first among open-weight models anywhere on that leaderboard.
| Model | SWE-bench Verified | Measured by |
|---|---|---|
| Claude Opus 5 | 97.00% | Vals AI |
| DeepSeek V4-Pro 0813 | 96.40% | Vals AI |
| Kimi K3 | 93.40% | Vals AI |
| Claude Opus 4.8 | 88.60% | Vals AI |
| Grok 4.5 | 86.60% | Vals AI |
Vals runs every system on its board through one harness: a minimal bash-tool-only agent, 500 real GitHub issues in Docker containers, scored by whether the generated patch passes the repository's own unit tests. No model gets a bespoke scaffold. Vals lists V4-Pro 0813's full run at $0.02 per test.
What makes this result unusually credible is what DeepSeek didn't do: the company never published a SWE-bench figure of its own. There is no vendor chart here to discount. The number is purely third-party, on a neutral rig, and it is 0.6 points off the best model in the world.
DeepSeek V4-Pro 0813 does not appear on the Aider Polyglot leaderboard at all, so there is currently no multi-language edit-format score to compare against this.
What does DeepSeek V4-Pro score on LiveBench?
Its LiveBench global average is 77.44, and its agentic coding score is 54.95. We computed both from LiveBench's own published task-level results table, retrieved on 23 August 2026 from livebench.ai, using LiveBench's own category definitions.
| Model | Global average | Agentic coding |
|---|---|---|
| Claude Fable 5 (max effort) | 82.97 | 62.17 |
| GPT-5.6 Sol (max) | 81.05 | 56.21 |
| Claude Opus 5 (max effort) | 80.08 | 65.20 |
| Kimi K3 | 79.19 | 62.17 |
| Qwen3.8-Max | 78.46 | 64.65 |
| Grok 4.6 | 78.04 | 57.02 |
| DeepSeek V4-Pro 0813 | 77.44 | 54.95 |
| DeepSeek V4-Flash 0731 | 74.17 | 46.77 |
| DeepSeek V4-Pro (April preview) | 71.57 | 42.63 |
Against those seven frontier peers, V4-Pro 0813 is last on agentic coding. That is the honest headline, and it is the mirror image of the SWE-bench result.
Two details make the number more useful than the ranking alone. First, LiveBench's "agentic coding" category is three task families — JavaScript, TypeScript and Python — and V4-Pro's sub-scores are wildly uneven: JavaScript 68.18, Python 50.00, TypeScript 46.67. TypeScript is where it collapses. It is worth noting that Claude Opus 5 scores 43.33 on the same TypeScript tasks, below DeepSeek; Opus 5 wins the category on JavaScript (77.27) and Python (75.00). So "DeepSeek is bad at agentic coding" is really "DeepSeek is bad at agentic Python", which is a much more actionable thing to know.
Second, the GA checkpoint is a genuine improvement over the preview and not a relabel. The April deepseek-v4-pro entry sits at 71.57 global / 42.63 agentic. The 0813 checkpoint gains +5.9 global and +12.3 on agentic coding — an independent corroboration of DeepSeek's own claim that GA was primarily an agent upgrade.
Why do three sources report three different Terminal-Bench scores?
Because the harness is doing most of the work, and this is the single most misread figure attached to the model.
| Source | Harness | Terminal-Bench 2.1 |
|---|---|---|
| DeepSeek (vendor-reported) | DeepSeek's own harness | 87.9 |
| Artificial Analysis | AA's Intelligence Index rig | 79% |
| Vals AI | Terminus 2 reference harness | 54.68% (±1.50), rank 36/57 |
The 87.9 comes from DeepSeek's 13 August changelog entry and the 0813 model card. The 54.68% comes from Vals running the benchmark's own reference harness. A 33-point spread is far outside the 7–13 points models typically lose when moved onto a neutral scaffold, and Artificial Analysis landing at 79% in the middle suggests the truth depends heavily on how much scaffolding you are willing to build.
The tiebreaker: the official Terminal-Bench 2.1 leaderboard currently lists 17 verified entries — topped by Claude Code with Fable 5 at 83.8% — and no DeepSeek V4 model appears on it at all. Until DeepSeek submits, treat 87.9 as a claim about DeepSeek's harness, not about the model.
What does Artificial Analysis measure it at?
Artificial Analysis puts DeepSeek V4-Pro 0813 at 53 on its Intelligence Index, a composite of ten evaluations. Claude Opus 5 scores 63 and Claude Fable 5 scores 62. V4-Flash 0731 scores 52 — meaning the flagship buys you exactly one index point over its own budget sibling.
The component scores show where the ten points to the frontier go. GPQA Diamond is effectively saturated at 93% — graduate-level reasoning is not the problem. The losses are in multi-step tool and domain work: GDPval-AA v2 55%, SciCode 49%, τ³-Banking 40%, Humanity's Last Exam 39% without tools, and CritPt 18%, the weakest component by a distance.
AA-Omniscience accuracy is 49%, against Claude Opus 5's 61% and Fable 5's 65%. That one deserves care, because AA's headline Omniscience Index is not the accuracy figure — it runs from −100 to 100, rewards correct answers, penalises hallucinations, and applies no penalty for declining to answer. Fable 5 leads that index at 43 and Opus 5 at 37. Accuracy tells you what V4-Pro knows; it says nothing about how often it invents an answer rather than abstaining, and in production that distinction matters more.
On speed the trade runs the other way. AA measures V4-Pro 0813 at 76.8 output tokens/second with 1.74s time-to-first-token, against Claude Opus 5's 56 tokens/second and 32.30s TTFT. For interactive work that is not a rounding error.
What is DeepSeek V4-Pro's knowledge cutoff?
DeepSeek does not publish one. That is the accurate answer, and it is worth stating plainly because several aggregators present a date as though it were official.
We checked the 0813 model card, the API changelog, and DeepSeek's API documentation. None of them state a training-data cutoff. What is actually known:
- An upper bound exists. Pretraining necessarily concluded before the V4 paper (arXiv 2606.19348) and the 24 April 2026 preview. The 0813 checkpoint is a post-trained and speculative-decoding-augmented version of that same base, not a re-pretrain, so August 2026 knowledge should not be assumed.
- Third-party "Apr 2026" listings are inferred, not sourced. Aggregators that publish a cutoff date list April 2026, which is simply the preview release date. No DeepSeek document supports it.
- Asking the model is unreliable. GitHub issue deepseek-ai/DeepSeek-V3#1389, opened 2 June 2026 and still open with no maintainer reply, documents Expert Mode — which runs V4-Pro — self-reporting as "DeepSeek-V3.2 with May 2025 knowledge cutoff". Instant Mode in the same app correctly identifies itself as V4-Flash. Self-reported cutoffs come from the system prompt and post-training data, not from the weights.
How to establish it yourself in ten minutes: open a fresh session with web search disabled, then ask about datable events at monthly granularity across a sliding window — say October 2025 through May 2026 — in three unrelated domains (a software release, a sports result, a public-market event). Recall will be solid, then patchy, then absent. The month where accuracy falls off across all three domains is your practical cutoff for that model. Do not accept the model's own stated date as evidence either way.
What actually shipped in the GA release?
From DeepSeek's changelog entry dated 13 August 2026:
- Reasoning-effort control. Three thinking-effort levels —
low,high,max— exposed via thereasoning_effortparameter. DeepSeek recommends allowing 384K output tokens at high and max. - Native OpenAI Responses API support, described in the changelog as "specifically adapted for Codex" — so Codex-style tooling points at DeepSeek without a translation layer.
- Agent-capability upgrades, rolled out simultaneously across the app, web (Expert Mode) and API.
- Model alias unchanged.
deepseek-v4-pronow resolves to the 0813 checkpoint, so existing integrations were upgraded silently. - DSpark shipped in the weights. vLLM enables it with
--speculative-config; SGLang with--speculative-algorithm DSPARK. Recommended sampling is temperature 1.0, with top_p 0.95 for agentic use and 1.0 otherwise.
The model card's own preview-to-GA deltas are large and consistent with the LiveBench movement we measured: Terminal-Bench 2.1 +15.8, DeepSWE +50.0, Cybergym +30.6, DSBench-FullStack +29.3, NL2Repo +23.0, HLE-with-tools +11.8. Vendor-reported, but the direction is independently corroborated.
One thing that did not ship: vision. V4-Pro remains text-only. DeepSeek's 21 August release added deepseek-v4-flash-vision-exp, an experimental multimodal build of Flash — not Pro.
How much does a benchmark point cost?
This page is not the rate card — full tables and worked examples live in our DeepSeek V4-Pro pricing reference, and the scheduling detail in the peak/off-peak billing guide. But cost is what makes a benchmark number mean anything, so here is the minimum.
Since 16 August 2026 V4-Pro bills at $0.66 / $1.98 per million input/output off-peak, doubling to $1.32 / $3.96 at peak (01:00–04:00 and 06:00–10:00 UTC, weekdays). Artificial Analysis blends that to roughly $0.69 per million tokens against $3.85 for Claude Opus 5.
So the trade on patch-shaped work is: give up 0.6 points of SWE-bench Verified, pay about one-fifth the blended rate, and get first-token latency roughly 18x faster. On agentic work the trade inverts and it is not close — 10.25 LiveBench agentic points behind Opus 5 is not something a price advantage compensates for when the agent has to finish unattended.
Can you run DeepSeek V4-Pro locally?
Realistically, no. The 0813 weights are 892.8 GB across 92 files. They are MIT-licensed and freely downloadable, and vLLM and SGLang both support the DSpark path — but the hardware required puts this outside anything most teams own.
V4-Flash 0731 is the practical open-weight option at 166.9 GB with 284B total and 13B activated parameters, which is exactly why it has 55x the download count. If self-hosting is the goal, plan around Flash and read our DeepSeek V4 VRAM and GPU requirements breakdown before buying anything.
Is DeepSeek V4-Pro 0813 good at agentic coding?
Not relative to its price-peers. Three independent lines of evidence point the same way: last of seven frontier models on LiveBench agentic coding (54.95), 54.68% on Terminal-Bench 2.1 under the reference harness against a claimed 87.9, and absent from the official Terminal-Bench leaderboard entirely.
The nuance that saves it: those are all long-horizon autonomy measurements. Vals' SWE-bench harness is agentic too — the model gets bash and has to navigate a real repository — and there it comes second in the world. The distinction is duration and error recovery, not tool use. V4-Pro is excellent on bounded, verifiable tasks and unreliable when nobody checks its work for an hour.
If you are choosing between open-weight flagships for agent work specifically, Kimi K3 is the comparison that matters. It loses to V4-Pro on SWE-bench Verified (93.40% against 96.40%) but beats it clearly on LiveBench agentic coding (62.17 against 54.95) and on global average (79.19 against 77.44). Bounded work favours DeepSeek; open-ended work favours Kimi. We covered that matchup in DeepSeek V4 vs Kimi K3.
Should you use DeepSeek V4-Pro 0813?
A decision rule rather than a verdict:
- Yes, for patch-shaped work. Bug fixes, targeted changes, PR-scale edits with tests to verify against. Second in the world on SWE-bench Verified at $0.02 per test, measured by someone with no stake in the result, is not a marginal offer.
- Yes, when latency matters. 1.74s to first token against Opus 5's 32.30s changes what kind of product you can build on it.
- No, for unattended long-horizon agents. Three independent measurements agree, and the one that disagrees is DeepSeek's own unreleased harness.
- No, for TypeScript-heavy agent loops specifically. 46.67 on LiveBench's TypeScript tasks is the sharpest single weakness in the profile.
- Check, if factual reliability is critical. 49% AA-Omniscience accuracy against Opus 5's 61% is a real gap, and DeepSeek's abstention behaviour is not published.
The broader read: DeepSeek shipped an agent-tuning and serving-efficiency release that closed a lot of ground on its own preview without closing the gap to Anthropic on autonomy. If you already know which of your tasks are bounded, that is a very good deal. If you were hoping to point an agent at a repo and walk away, it is not yet the model for it. For wider family context — Pro versus Flash, the architecture, the API surface — start with our DeepSeek V4 complete guide.
FAQ
What is DeepSeek V4-Pro 0813?
It is the general-availability checkpoint of DeepSeek V4-Pro, published 13 August 2026 and superseding the 24 April preview. It is a 1.6T-parameter mixture-of-experts model with 49B activated parameters, a 1M-token context window, MIT-licensed weights totalling 892.8 GB, and the DSpark speculative-decoding module built in.
What does DeepSeek V4-Pro score on SWE-bench Verified?
96.40%, ranking second of 82 models on Vals AI's leaderboard as of 19 August 2026 and first among open-weight models. Only Claude Opus 5 scores higher, at 97.00%. Vals uses one neutral bash-tool-only harness for every model and recorded the run at $0.02 per test. DeepSeek itself never published a SWE-bench figure.
What does DeepSeek V4-Pro score on LiveBench?
77.44 global average and 54.95 on agentic coding, computed from LiveBench's published results table on 23 August 2026. That agentic score is last among seven frontier peers — Claude Opus 5 leads at 65.20. Sub-scores are JavaScript 68.18, Python 50.00 and TypeScript 46.67, so the weakness is concentrated rather than uniform.
What is DeepSeek V4-Pro's knowledge cutoff?
DeepSeek does not publish one. The model card, API docs and changelog are all silent on training-data cutoff. Third-party listings of "April 2026" appear to be inferred from the preview release date. Asking the model is unreliable — an open GitHub issue documents Expert Mode self-reporting a May 2025 cutoff and the wrong model version entirely.
Can you run DeepSeek V4-Pro locally?
Not on hardware most teams have. The 0813 weights are 892.8 GB across 92 safetensors files. They are MIT-licensed and vLLM and SGLang both support them, but the memory requirement is prohibitive. V4-Flash 0731 at 166.9 GB, with 284B total and 13B activated parameters, is the realistic self-hosting choice and has roughly 55x the downloads.
Is DeepSeek V4-Pro good at agentic coding?
For bounded agentic tasks, yes — it is second in the world on SWE-bench Verified, which is itself a bash-agent benchmark. For long-horizon autonomous work, no: it is last of seven peers on LiveBench agentic coding, scores 54.68% on Terminal-Bench 2.1 under the reference harness against a claimed 87.9, and does not appear on the official Terminal-Bench leaderboard.
How much does DeepSeek V4-Pro cost?
Since 16 August 2026 it bills at $0.66 per million input and $1.98 per million output off-peak, doubling to $1.32 and $3.96 during peak hours of 01:00–04:00 and 06:00–10:00 UTC on weekdays. Cache reads are $0.022 off-peak. Full rate tables and worked examples are in our dedicated DeepSeek V4-Pro pricing reference.
Is DeepSeek V4-Pro better than V4-Flash?
On measured coding ability, yes: 77.44 against 74.17 on LiveBench global average, and 54.95 against 46.77 on agentic coding. But on Artificial Analysis's Intelligence Index the gap is a single point — 53 against 52. Flash is also the only V4 model that is practical to self-host at 166.9 GB, and the only one with a vision variant. For many workloads the flagship premium is hard to justify.
Does DeepSeek V4-Pro support images?
No. V4-Pro 0813 is text-only. DeepSeek released an experimental multimodal model on 21 August 2026 as deepseek-v4-flash-vision-exp, which tokenises images at up to 384 tokens each and bills at V4-Flash rates. It is built on Flash, not Pro, and is marked experimental.