Grok 4.6: Release Date and What Changed (August 2026)
Quick answer. Grok 4.6 launched on August 12, 2026, 35 days after Grok 4.5. It keeps the 500K context window and $2/$6 per million token pricing, adds an xhigh reasoning level, and lifts the Artificial Analysis Intelligence Index from 56 to 61. It reached the API, Cursor, GitHub Copilot, Bedrock and Grok Build through August.
xAI shipped Grok 4.6 on August 12, 2026, roughly five weeks after Grok 4.5. The API model id is simply grok-4.6, and xAI's own framing for the release is that it was built "for long-running agents and more ambitious interactive and visual work."
The launch-day story was narrow — API, Cursor and Grok Build. The three weeks since have been the more interesting part: the model rolled out across GitHub Copilot, Amazon Bedrock, Google's Model Garden and Microsoft Foundry, and Grok Build went from an invite-shaped tool to a public product on every plan. This page covers the release itself, what actually changed, what the benchmark numbers mean, and where you can use it as of the end of August 2026.
When was Grok 4.6 released?
August 12, 2026. That is the date the model went live on the xAI API and in Cursor and Grok Build simultaneously. But "released" has meant something slightly different every week since, because availability kept expanding. Here is the full dated timeline.
| Date (2026) | What shipped |
|---|---|
| August 11 | Grok Bot enters early beta — persistent cloud "AI teammates" for SuperGrok Heavy, Cursor Ultra and Cursor Teams Premium, on desktop and iOS |
| August 12 | Grok 4.6 launches. Live on the xAI API, in Cursor and in Grok Build, with 2x included usage for the first week. Gateway partners: OpenRouter, Vercel, Cloudflare |
| August 14 | Grok 4.6 appears in GitHub Copilot's model picker (VS Code and GitHub workflows) |
| August 15 | Grok Build 1.0.5 — environment variable overrides, stability fixes |
| August 19 | Grok Build goes public on all plans, web and mobile, with publishing, sharing, X integration, custom domains, GitHub export, secrets and connectors. Grok 4.6 reaches general availability on Amazon Bedrock |
| August 21 | Grok 4.6 added to Google's Model Garden / Enterprise Agent Platform. Grok Bot access widens to SuperGrok Plus, SuperGrok Heavy, Cursor Pro+, Cursor Ultra and Cursor Teams |
| August 26 | Grok 4.6 available via Microsoft Foundry |
If you were told in mid-August that Grok 4.6 was an API-only, developer-only release, that was true for about 48 hours. By August 21 it was reachable from all three major clouds and from consumer subscription tiers.
What are Grok 4.6's specifications?
| Spec | Grok 4.6 |
|---|---|
| Released | August 12, 2026 |
| API model id | grok-4.6 |
| Context window | 500,000 tokens |
| Modality | Text and image in, text out |
| Knowledge cutoff | February 1, 2026 |
| Reasoning levels | low / medium / high (default) / xhigh (new) |
| Regions | us-east-1, us-west-2 — no EU region |
| Rate limits | 150 requests/second, 50M tokens/minute |
| Batch API | Not supported |
| Tool support | Function calling and structured outputs |
xAI has published no architecture details for Grok 4.6 — no parameter count, no mixture-of-experts disclosure, no training compute figure and no system card. Any parameter number you see quoted for this model is third-party inference, not a vendor statement.
Is there a separate "Grok 4.6 High" model?
No, and this trips people up constantly. You will see "Grok 4.6 (high)" and "Grok 4.6 (xhigh)" listed as separate rows on benchmark leaderboards. Those are reasoning_effort settings against one model id, not different SKUs. There is exactly one grok-4.6.
The genuinely new setting is xhigh — "maximum reasoning depth, with correspondingly higher latency." It is the first release where that value does anything: Grok 4.5 accepted the parameter and silently downgraded the request to high. Worth knowing before you reach for it: on Artificial Analysis's index, xhigh currently scores one point lower than the default high. More thinking is not automatically better.
What actually changed in Grok 4.6?
xAI describes Grok 4.6 as a post-training and training-recipe release rather than a new foundation. The changes it names are:
- Longer supplemental training on curated model-generated data aimed at reasoning
- An improved optimizer and training recipe
- Better self-testing and verification on extended, long-horizon tasks
- Stronger visual and interactive work — the "ambitious interactive and visual projects" framing
- The
xhighreasoning level, the one concrete API-surface addition
What did not change: the context window (still 500K), the price list ($2/$6), the modalities (text and image in, text out), and the region footprint. If you are already calling grok-4.5, swapping the model string is the entire migration.
How is Grok 4.6 different from Grok 4.5 in measured terms?
Vendor descriptions are one thing; independent measurement is another. Artificial Analysis has run both models under the same harness, which makes this the cleanest generation-over-generation comparison available.
Metric (Artificial Analysis, default high) | Grok 4.5 | Grok 4.6 |
|---|---|---|
| AA Intelligence Index | 56 | 61 |
| Output speed | 54.2 tok/s | 59.2 tok/s |
| Time to first token | 14.51s | 43.82s |
| Blended price per 1M tokens | $1.21 | $1.35 |
| Context window | 500K | 500K |
Two things stand out. The five-point intelligence gain is real and large for a point release. And the time to first token tripled — from about 14.5 seconds to nearly 44. That is the single biggest practical regression in the release, and it is a direct consequence of the deeper reasoning the model now does before it starts emitting. Throughput once it starts is actually slightly better than 4.5.
Note also that blended cost rose about 12% despite an unchanged price list. That is the model spending more reasoning tokens per task. Per-token pricing held; cost per job did not.
For a section-by-section breakdown of the two models, see Grok 4.6 vs Grok 4.5: what changed and should you upgrade.
What do Grok 4.6's benchmark numbers actually say?
Separate these into two buckets, because they are not the same kind of claim.
What xAI reported
These are the figures in xAI's own launch announcement, all for Grok 4.6 at the default high reasoning level:
| Benchmark | Grok 4.6 (xAI-reported) |
|---|---|
| AA Intelligence Index | 61 |
| GDPVal-AA v2 | 1753 |
| CursorBench v3.2 | 69.9% |
| DeepSWE v1.1 | 65.9% |
| FrontierCode v1.1 (Extended) | 61.3% |
| APEX-Agents | 57.5% |
| APEX-SWE | 56.4% |
| AA-Briefcase | 1577 |
| Terminal-Bench v3.0 | 26% |
| Harvey LAB (Vals) | 15.8% |
What independent evaluators measured
Artificial Analysis confirms the headline: Grok 4.6 at high scores 61 on its Intelligence Index, placing it 6th of 187 models. That is a clean result — the vendor's launch number survived contact with the independent evaluator, which is not always the case. AA also measures output speed at 59.2 tokens/second (below the median for reasoning models), time to first token at 43.82 seconds, and a total cost of $1,157.64 to run its full index against the model.
Arena (the human-preference leaderboard formerly at LMArena) tells a very different story. In its August 27, 2026 snapshot, Grok 4.6 sits at rank 48 with a preliminary score of 1461 ±10, well outside a top ten dominated by Anthropic's Claude family and Meta's Muse Spark. Preliminary scores move as votes accumulate, so treat that as provisional — but the gap between "6th on a benchmark index" and "48th on blind human preference" is too wide to ignore.
The Terminal-Bench version trap
This is the single most misread number in the release. xAI reports 26% on Terminal-Bench v3.0. Most third-party leaderboards still report v2.1, a substantially easier version, where scores in the 70–90% range are normal.
Those are the same model on different exams. Any article, chart or vendor comparison that places a v3.0 figure next to a v2.1 figure is producing nonsense. Always check the version suffix before comparing Terminal-Bench scores across sources. We break the cross-source discrepancies down in detail in Grok 4.6 benchmarks explained.
Where can you actually use Grok 4.6?
As of August 31, 2026, five routes:
- The xAI API —
grok-4.6via console.x.ai. Two regions (us-east-1, us-west-2), 150 requests/second, 50M tokens/minute, no Batch API support. - The major clouds — generally available on Amazon Bedrock (August 19), Google's Model Garden / Enterprise Agent Platform (August 21) and Microsoft Foundry (August 26). All three list the same $2/$6 headline rate.
- Gateways — OpenRouter, Vercel and Cloudflare were day-one partners.
- Coding tools — Cursor from day one; GitHub Copilot's model picker from August 14, in both VS Code and GitHub workflows.
- xAI's own products — Grok Build, now public on all plans across web and mobile since August 19. Grok Bot, the persistent cloud "AI teammate" product, is available on SuperGrok Plus, SuperGrok Heavy, Cursor Pro+, Cursor Ultra and Cursor Teams since the August 21 expansion.
Grok Build is the surface most people underestimate — the August 19 public release added publishing, custom domains, GitHub export, secrets and connectors, which turns it from a demo into something you can actually ship from. If that is the part you care about, we have a dedicated walkthrough: the xAI Grok Build skills and connectors guide.
The EU gap is still open. Only us-east-1 and us-west-2 are offered on the xAI API. If your data residency requirements rule out US-only inference, the cloud marketplace routes are worth checking before you rule the model out entirely.
What does Grok 4.6 cost?
The short version: $2 per million input tokens and $6 per million output, with cached input at $0.50 — identical to Grok 4.5's list price. Two details that matter more than the headline:
- Crossing 200K prompt tokens doubles the rate on the entire request. Above that threshold it is $4 input / $1 cached / $12 output, applied to every token in the call, not just the ones past the line. A 210K-token prompt is billed at the higher rate across all 210K.
- xAI's
x_searchtool costs $5 per 1,000 calls — first-party licensed access to X posts, profiles and threads. Nothing else on the market offers it.
That is deliberately a summary. Cache-hit economics, per-task cost against competitors, the long-context threshold maths and worked agent-loop examples all live in our dedicated breakdown: Grok 4.6 pricing and API costs. If cost is the reason you are here, start there instead.
Where does Grok 4.6 sit against the frontier?
Honestly: frontier-adjacent, not frontier-leading, and the answer depends heavily on which measurement you trust.
On Artificial Analysis's index, the top of the board looks like this:
| # | Model | AA Intelligence Index |
|---|---|---|
| 1 | Claude Opus 5 (max) | 63 |
| 2 | Claude Opus 5 (xhigh) | 63 |
| 3 | Claude Fable 5 | 62 |
| 4 | Claude Opus 5 (high) | 61 |
| 5 | GPT-5.6 Sol (max) | 61 |
| 6 | Grok 4.6 (high) | 61 |
| 7 | Grok 4.6 (xhigh) | 60 |
| 8 | Kimi K3 (max) | 60 |
Read that carefully. Grok 4.6 is tied on the rounded score with Claude Opus 5 (high) and GPT-5.6 Sol (max) — a genuine achievement at $2/$6 against considerably more expensive models. But it is behind the actual leaders, and the Arena result suggests the benchmark placement flatters it relative to how people rate its answers in blind comparison.
Two things are unambiguously in its favour. The price-to-index ratio is excellent: nothing else at $2/$6 scores 61. And real-time X data via x_search is not substitutable — for social listening, news monitoring or anything needing live public sentiment, no competitor has licensed firehose access.
Two things are unambiguously against it. The 500K context window is last in class among current frontier models, most of which are at or above 1M — and the 200K billing threshold means you pay double well before you exhaust it. And a 43-second time to first token is difficult to build interactive tooling around.
Should you use Grok 4.6?
Yes, if you are running batch or background agent work where a 43-second cold start is irrelevant and the index-per-dollar ratio is what matters. This is the model's strongest case.
Yes, if you need live X data. The x_search tool has no equivalent anywhere else.
Yes, if you are already on Grok 4.5. The upgrade is a model-string change for five index points, and there is no reason to stay.
Probably not, if latency is user-visible. Nearly 44 seconds to first token will be felt in any interactive surface.
Probably not, if you routinely exceed 200K tokens of context. You will hit the doubled rate long before the 500K ceiling, and competitors offer 1M windows without the cliff.
Probably not, if output quality as humans judge it is your bar rather than benchmark placement. The Arena gap is the number to weigh there.
Not yet, if you require EU data residency on the xAI API directly.
The practical decision rule most teams land on: use Grok 4.6 as the cheap, capable implementer, not as the model you reach for first. Plan with whatever sits at the top of your quality bar, hand the mechanical execution to Grok, and let the price difference compound across the volume. That is a more defensible position than any frontier claim, and it is one Grok 4.6 holds comfortably.
FAQ
When was Grok 4.6 released?
August 12, 2026 — about 35 days after Grok 4.5. It launched simultaneously on the xAI API, in Cursor and in Grok Build, with gateway availability on OpenRouter, Vercel and Cloudflare from day one. Cloud availability followed through the rest of August: Amazon Bedrock on August 19, Google's Model Garden on August 21, and Microsoft Foundry on August 26.
What's new in Grok 4.6?
xAI describes it as a post-training release: longer supplemental training on curated model-generated data, an improved optimizer and training recipe, and better self-verification on long-horizon tasks. The one concrete API addition is the xhigh reasoning effort level. Context, price, modalities and regions are all unchanged from Grok 4.5.
How is Grok 4.6 different from Grok 4.5?
Measured by Artificial Analysis under the same harness, Grok 4.6 scores 61 on the Intelligence Index against Grok 4.5's 56, and generates 59.2 tokens/second against 54.2. The trade-off is latency: time to first token rose from 14.51s to 43.82s, and blended cost per million tokens rose from $1.21 to $1.35. Same 500K context on both.
How much does Grok 4.6 cost?
$2 per million input tokens, $0.50 cached input, $6 per million output — unchanged from Grok 4.5. Prompts of 200K tokens or more are billed at $4/$1/$12 across the entire request, not just the excess. Full breakdown, including cache economics and per-task comparisons, is in our Grok 4.6 pricing and API costs guide.
What is Grok 4.6's context window?
500,000 tokens — unchanged from Grok 4.5, and the smallest among current frontier models, most of which offer 1M or more. Practically it is smaller still, because prompts at or above 200K tokens are billed at double the standard rate across the whole request.
Is Grok 4.6 better than Claude Opus 5?
Not on overall capability. Claude Opus 5 leads the Artificial Analysis Intelligence Index at 63 against Grok 4.6's 61, and Anthropic's models dominate the top of Arena's human-preference leaderboard while Grok 4.6 sits at rank 48. Grok 4.6's advantage is price: $2/$6 against Opus 5's considerably higher rate for a two-point index gap.
Is there a "Grok 4.6 High" model?
No. "High" is the default reasoning_effort setting on the single grok-4.6 model id, not a separate SKU. The four supported levels are low, medium, high and xhigh. Benchmark sites list them as separate rows, which is what causes the confusion.
Should I use xhigh reasoning effort?
Usually not by default. xhigh is new in Grok 4.6 — Grok 4.5 silently downgraded it to high — and it delivers maximum reasoning depth at correspondingly higher latency. But on Artificial Analysis's index it currently scores one point below the default high setting. Reserve it for genuinely hard problems where you can measure the difference.
Is Grok 4.6 available in the EU?
Not on the xAI API directly — only us-east-1 and us-west-2 regions are offered. If EU data residency is a requirement, check the Amazon Bedrock, Google Model Garden and Microsoft Foundry routes, which have their own regional footprints, before ruling the model out.