Quick answer. Claude Opus 5 leads or ties the frontier across most 2026 benchmarks: 43.3% on Frontier-Bench (agentic coding), 30.2% on ARC-AGI-3 (novel reasoning, roughly 3x the next model), and a 1,861 GDPval-AA v2 Elo for knowledge work. It ships at $5 / $25 per million tokens, half the input price of Claude Fable 5. The real story is capability per dollar.
Anthropic launched Claude Opus 5 on July 24, 2026, and framed it plainly: a thoughtful, proactive model that comes close to the frontier intelligence of the larger Claude Fable 5 at roughly half the price. It is the fourth Claude model in under two months, and it is pitched as the everyday workhorse for coding, agentic tasks, and knowledge work rather than a specialist flagship.
This post is a single, complete reference for what every benchmark Anthropic reported actually measures, Opus 5's exact score on each, and the comparison models it was tested against. If you want the full product picture — availability, the effort toggle, fast mode, migration notes — read the Claude Opus 5 launch guide. Here we stay on the numbers.
What benchmarks did Anthropic report for Claude Opus 5?
Anthropic and independent trackers published results across seven public evaluations plus science and security suites. Here is the master table, with the strongest reported comparison model on each.
| Benchmark | What it measures | Opus 5 | Best comparison |
|---|---|---|---|
| Frontier-Bench v0.1 | Agentic coding (% of tasks passed) | 43.3% | GPT-5.6 Sol 34.4% · Fable 5 33.7% |
| ARC-AGI-3 | Novel problem-solving, no memorization | 30.2% | GPT-5.6 Sol 7.8% |
| GDPval-AA v2 | Human-graded economic knowledge work (Elo) | 1,861 | Fable 5 1,747 · GPT-5.6 Sol 1,736 |
| SWE-bench Pro | Real GitHub issues resolved | 79.2% | Mythos 5 80.3% · Fable 5 80.0% |
| CursorBench 3.2 | In-editor coding at max effort | Within 0.5 pts of Fable 5's peak | Fable 5 (peak) |
| Zapier AutomationBench | End-to-end automation workflows | 100% (full churn-prevention flow) | ~1.5x throughput of next-closest |
| OSWorld 2.0 | Computer use (screen + apps) | Beats Fable 5's peak at ~1/3 the budget | Fable 5 (peak) |
Two patterns jump out. Opus 5 wins outright on the hardest reasoning and agentic-coding evals, and it either matches or edges the frontier on everything else while costing meaningfully less to run. Below, each benchmark in turn.
How does Opus 5 do on agentic coding (Frontier-Bench)?
Frontier-Bench v0.1 scores agentic coding — a model working through multi-step engineering tasks the way an autonomous coding agent would, measured as the percentage of tasks it passes end to end. Opus 5 posted 43.3%, ahead of GPT-5.6 Sol at 34.4% and Claude Fable 5 at 33.7%. Anthropic notes it more than doubles Opus 4.8's score on this eval at a lower cost per task, which is the sharpest single data point on the generational jump. This is one of the few benchmarks where Opus 5 beats the larger Fable 5 outright.
How good is Opus 5 at novel reasoning (ARC-AGI-3)?
ARC-AGI-3 is designed to resist memorization: it presents genuinely novel problems so a model can't lean on patterns absorbed during training. It's the closest public proxy for reasoning that generalizes. Opus 5 scored 30.2% versus 7.8% for GPT-5.6 Sol — roughly three times the next-best publicly listed model. Fable 5's number wasn't published, but the gap over GPT-5.6 Sol is the widest margin Opus 5 shows on any benchmark, and it's the result that most supports Anthropic's "comes close to the frontier" framing.
Can Opus 5 handle real economic knowledge work (GDPval-AA v2)?
GDPval-AA v2 is human-graded and expressed as an Elo rating rather than a percentage. It measures performance on economically valuable knowledge work — the kind of drafting, analysis, and synthesis a knowledge worker is actually paid for — with human raters comparing outputs head to head. Opus 5 reached 1,861, ahead of Fable 5 (1,747) and GPT-5.6 Sol (1,736). For teams evaluating a model as a day-to-day assistant rather than a coding engine, this is the eval that maps most directly to real work.
Where does Opus 5 land on SWE-bench Pro and CursorBench?
SWE-bench Pro tests resolution of real GitHub issues — patch a live repository so the fix passes the project's own tests. Opus 5 scored 79.2%, third behind Mythos 5 (80.3%) and Fable 5 (80.0%), but the headline is the leap over its predecessor: Opus 4.8 sat at 69.2%, so Opus 5 closes most of the gap to the leaders while sitting under a single point behind them.
CursorBench 3.2 measures in-editor coding — the assist-while-you-type workflow developers live in. At max effort, Opus 5 lands within 0.5 points of Fable 5's peak at roughly half the cost per task. If your team already runs AI in the editor, that half-price-for-near-parity trade is the practical takeaway. For the wider landscape of tools this plugs into, see our guide to AI coding agents.
What about agent workflows and computer use?
Zapier AutomationBench checks whether a model can drive a full automation workflow to completion. Opus 5 completed an end-to-end churn-prevention workflow at 100%, and Anthropic reports it runs about 1.5x the throughput of the next-closest model at matching cost.
OSWorld 2.0 is the computer-use benchmark: the model operates a real screen — clicking, typing, navigating apps — to finish tasks. Anthropic says Opus 5 "outperforms rivals at every price point" and beats Fable 5's peak at roughly one-third of the budget. Together these two results are the clearest evidence that Opus 5's value isn't just raw accuracy — it's accuracy per dollar on the long-running, tool-using tasks that define agentic work.
How does Opus 5 do on science and security?
On life sciences, Opus 5 improves on Opus 4.8 across every eval, with organic chemistry up about 10 points. Anthropic calls it its "most capable generally available model for scientific research," with strong biology performance.
On cybersecurity (OSS-Fuzz), Opus 5 tracks close to Mythos 5 on vulnerability discovery but sits behind it on exploitation. It can examine source code but won't scan compiled binaries — a deliberate boundary. For the hardest frontier cybersecurity exploitation and the toughest biology research, Anthropic still points to Mythos 5.
What does "capability per dollar" actually mean here?
The through-line across every benchmark is price. Opus 5 holds the same pricing as Opus 4.8 while delivering near-frontier results, which is why "capability per dollar" — not any single leaderboard win — is the framing Anthropic leans on.
| Model / mode | Input ($/M tokens) | Output ($/M tokens) | Notes |
|---|---|---|---|
| Opus 5 (standard) | $5 | $25 | Same price as Opus 4.8 |
| Opus 5 (fast mode) | $10 | $50 | 2x price, ~2.5x faster |
| Claude Fable 5 | $10 | — | 2x Opus 5's input price |
A few specifics reinforce the value angle. Opus 5 carries a 1M-token context window as both default and maximum, extended thinking is on by default, and a per-request effort toggle (low / medium / high) lets you dial reasoning compute up for hard problems or down for routine tasks — Anthropic's explicit lever to "balance cost and capability." Opus 5 is also not subject to the 30-day data-retention policy that applies to Fable 5, and its safety classifiers are expected to trigger about 85% less often, meaning fewer spurious refusals in day-to-day use. If you're weighing it against the previous workhorse, our Claude Opus 4.8 guide is the baseline to compare against.
Where does Opus 5 not lead?
It's an honest picture, not a clean sweep. Fable 5 and Mythos 5 still edge it on SWE-bench Pro, Anthropic recommends Fable 5 for the most advanced long-horizon autonomous work — agents running for days — and Mythos 5 remains ahead on frontier cybersecurity exploitation and the hardest biology research. Opus 5's pitch is that for the overwhelming majority of coding, agentic, and knowledge tasks, it delivers frontier-adjacent results at half the price, with the option to spend up via fast mode or high effort only when a task warrants it. For a direct frontier comparison, our GPT-5.6 vs Claude Fable 5 breakdown covers the tier above.
Benchmarks tell you what a model can do; shipping software tells you what it does under your constraints. If you're building an engineering team that can evaluate and deploy models like Opus 5 in real workflows, hire vetted remote developers through Codersera and extend your team with people who already work this way.
FAQ
What is Claude Opus 5's best benchmark result?
Its widest margin is on ARC-AGI-3, where it scored 30.2% versus 7.8% for GPT-5.6 Sol — roughly three times the next-best publicly listed model. It also wins Frontier-Bench agentic coding outright at 43.3%.
How much does Claude Opus 5 cost?
Standard pricing is $5 per million input tokens and $25 per million output tokens — the same as Opus 4.8 and half of Fable 5's input price. An optional fast mode doubles the price to $10 / $50 for roughly 2.5x the speed.
Is Claude Opus 5 better than Claude Fable 5?
It depends on the task. Opus 5 leads on Frontier-Bench, ARC-AGI-3, and GDPval-AA v2, while Fable 5 edges it on SWE-bench Pro and remains Anthropic's recommendation for the most advanced multi-day autonomous agent work. Opus 5's advantage is delivering near-frontier results at roughly half the price.
How does Opus 5 compare to GPT-5.6 Sol?
Opus 5 leads on every head-to-head reported: 43.3% vs 34.4% on Frontier-Bench, 30.2% vs 7.8% on ARC-AGI-3, and 1,861 vs 1,736 on GDPval-AA v2.
What does the effort toggle do?
The effort parameter (low / medium / high) controls how much reasoning compute Opus 5 spends per request. Low is faster and cheaper for routine tasks; high runs longer reasoning for hard problems. It's Anthropic's built-in lever to balance cost and capability without switching models.