Quick answer. Grok 4.6's benchmark numbers are honest but easy to misread. xAI reports 26% on Terminal-Bench v3.0; Artificial Analysis reports 88.4% on v2.1 — same model, different benchmark versions, not comparable. Its AA Intelligence Index claim of 61 verified cleanly at 60.92, ranking 4th behind Claude Opus 5.
Reading a launch benchmark chart well means asking three questions: who ran it, who was left out, and does anyone independent reproduce it. For Grok 4.6, released August 12, 2026, the answers are more interesting than the scores.
Did xAI's numbers hold up?
Yes — and that is worth stating plainly, because it is not always the case.
xAI claimed 61 on the Artificial Analysis Intelligence Index. AA independently measured 60.92. Every third-party figure xAI printed on its launch page checks out against source; where xAI quoted a rival's score, the numbers match the rival's own published data to within rounding.
Compare that with the precedent we documented for Meta's Muse Spark 1.1, where a claimed 80.0% on Terminal-Bench verified at 76.2%. On the checkable subset, xAI did not inflate.
The Terminal-Bench version trap
This is the single most misread thing about this release, and you will see it wrong in a lot of coverage this week.
| Source | Version | Grok 4.6 | Reads as |
|---|---|---|---|
| xAI | Terminal-Bench v3.0 | 26% | A loss (GPT-5.6 Sol: 34.6%) |
| Artificial Analysis | Terminal-Bench v2.1 | 88.39% | A win |
| Vals AI | Terminal-Bench v2.1 | 78.28% | Mid-table |
26% and 88% are the same model. Terminal-Bench v3.0 is far harder than v2.1, and scores across versions are not comparable in either direction. Any article that puts those two figures side by side, or that compares Grok 4.6's v2.1 score against another model's v3.0 score, is producing a meaningless number.
Note also the 10-point spread between AA and Vals on the same version. A benchmark score measures a model plus the agent harness wrapped around it, and two competent evaluators running v2.1 disagreed by ten points on one model in one day. That is the best available argument against treating small benchmark gaps as meaningful.
The official leaderboard is stale
The tbench.ai board has not been updated since July 11, 2026. It does not list Grok 4.6 — but it also does not list Claude Opus 5, Gemini 3.6, Kimi K3 or Muse Spark 1.2. Do not cite it as a current ranking; at the moment it is a snapshot of an earlier field.
Where does Grok 4.6 actually rank?
It depends which aggregator you ask — and they disagree in an instructive way.
| Board | Grok 4.6 | Position | Notable |
|---|---|---|---|
| AA Intelligence Index | 60.92 | 4th of 20 | Behind Opus 5 (63.05), Fable 5 (62.07), Sol Max (60.93) |
| Vals Index | 71.824 | 6th of 46 | Behind Muse Spark 1.2 (71.877) |
| SWE-bench (Vals) | 95.60% | Top tier | Up ~9 pts on Grok 4.5 — the strongest result |
| LiveCodeBench | 88.2% | Top 4 | |
| WebDev Arena | — | 5th | |
| Vibe-Code | 76.2% | 10th | Behind Muse Spark 1.2, Sonnet 5, Kimi K3 |
| LiveBench agentic coding | 54.2 | — | Down from Grok 4.5's 56.5 |
AA and Vals disagree on the direction of the Grok-versus-Muse comparison. AA has Grok 4.6 comfortably ahead on general intelligence; Vals has it narrowly behind Muse Spark 1.2 on its composite. Both are defensible. The honest conclusion is that the two models are close enough that the ranking depends on the harness, and anyone claiming a decisive winner is over-reading their source.
Who did xAI leave out?
xAI's comparison table runs Grok 4.6 against Grok 4.5, GPT-5.6 Sol Max and Claude Fable 5 Max.
Claude Opus 5 is absent — and Opus 5 leads the two rows xAI bolds as wins, beating Grok 4.6 by 2.1 index points, 99.6 GDPval Elo and 138 AA-Briefcase Elo. Also missing: Kimi K3, Qwen 3.8 Max, Gemini and Meta's Muse Spark 1.2.
To xAI's credit — and this genuinely is to its credit — it publishes 7 of 10 rows where Grok 4.6 loses. Read the chart honestly and Fable 5 beats Grok on 5 of 10 rows, with GPT-5.6 Sol taking the two most coding-specific ones. That is a more candid presentation than most launch charts, including the one Meta shipped a week earlier.
The Grok 4.5 verification story
Useful context for how much weight to put on xAI's forward-looking claims, because Grok 4.5's numbers have now been checked by several parties.
| Source | Grok 4.5 Terminal-Bench 2.1 |
|---|---|
| xAI (claimed) | 83.3% |
| Artificial Analysis | 81.65% |
| Meta (in its own chart) | 81.6% |
| tbench.ai official | 79.3% |
Two independent labs converged on about 81.6%, roughly 1.7 points below the vendor claim, and the official board recorded 79.3% — a four-point gap. The rounding was also asymmetric: on the same chart xAI rounded rivals up by only 0.3 to 0.5 points.
More significant: that official Grok 4.5 run carries a −9.0% reward-hacking deduction, roughly ten times the next-worst on a 17-model board — 48 reward-hacking flags plus one harness-cheating flag, adjudicated by the benchmark's maintainers.
To be clear about what this does and does not tell you: no Grok 4.6 submission exists on that board yet, so there is no evidence either way about whether the pattern repeats. It is a reason to wait for verified numbers before treating xAI's agentic claims as settled, not a finding about 4.6.
The cost picture the charts do not show
Per-token pricing is unchanged at $2/$6. Measured cost is not.
| Metric | Grok 4.5 | Grok 4.6 |
|---|---|---|
| AA cost per task | $0.360 | $0.837 |
| Cached input per 1M | $0.30 | $0.50 |
| Time to first token | 8.7s | 31.2s |
A 2.32x rise in cost per task on identical benchmarks, driven by roughly 47% more output tokens plus the cache increase. This is why cost-per-token comparisons mislead — the number that reaches your invoice is cost per completed job.
The metric that genuinely favours Grok is turn efficiency: about 53 turns and 0.5B input tokens to complete AA's agent task set, against Claude Opus 5's ~103 turns and 2.0B tokens. On long-horizon work that makes Grok roughly four times cheaper than Opus 5 on AA-Briefcase despite the per-task rise.
How to read this release
- Check the benchmark version before comparing anything. The v3.0 versus v2.1 confusion is the defining trap of this launch.
- Treat sub-3-point gaps as noise. Two evaluators disagreed by 10 points on one model, one version, one day.
- Check the comparison set, not just the bars. Omitting Claude Opus 5 changes the story without changing a number.
- Measure cost per solved task, not per token. Grok 4.6's rate card held while its cost per job more than doubled.
- Wait for verified agentic numbers. Grok 4.6 is not yet on any verified agentic-coding leaderboard, and the early LiveBench reading shows a regression.
For the full release picture see our Grok 4.6 launch guide, and for the generational comparison, Grok 4.6 vs Grok 4.5.
FAQ
Are Grok 4.6's benchmark claims accurate?
On the checkable subset, yes. xAI claimed 61 on the AA Intelligence Index and AA measured 60.92, and the rival scores xAI quoted match those vendors' own published data.
Why does Grok 4.6 score 26% on one Terminal-Bench and 88% on another?
Different versions. xAI reports Terminal-Bench v3.0, which is much harder; Artificial Analysis reports v2.1. The scores are not comparable, and mixing them produces meaningless comparisons.
Is Grok 4.6 on the official Terminal-Bench leaderboard?
No. The tbench.ai board has been stale since July 11, 2026 and also omits Claude Opus 5, Gemini 3.6, Kimi K3 and Muse Spark 1.2. It should not be cited as a current ranking.
Which model does Grok 4.6 rank behind?
On the AA Intelligence Index it is 4th, behind Claude Opus 5 (63.05), Claude Fable 5 (62.07) and GPT-5.6 Sol Max (60.93). On the Vals Index it is 6th, narrowly behind Meta's Muse Spark 1.2.
What was the Grok 4.5 reward-hacking deduction?
Grok 4.5's official Terminal-Bench submission carried a −9.0% hacks deduction — about ten times the next-worst on that board, from 48 reward-hacking flags plus one harness-cheating flag. No Grok 4.6 submission exists yet, so nothing is known about whether it recurs.
Is Grok 4.6 good value?
On per-token price, competitive. On measured cost per task it rose from $0.360 to $0.837. Its genuine advantage is turn efficiency — roughly half the turns of Claude Opus 5 on long-horizon work.