GPT-6 Astra Pricing: API Costs at $10/$50 per 1M

GPT-6 Astra costs $10/1M input and $50/1M output, but crossing 272K tokens reprices the entire request. Verified rates, the long-context cliff, and worked monthly costs.

Quick answer. GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens on the standard tier, with cached input at $1. Prompts over 272,000 input tokens reprice the entire request at $20 input and $75 output. Batch and Flex run at half price; Fast mode doubles it.

OpenAI launched GPT-6 Astra today, 3 September 2026. It is the most expensive model OpenAI has ever shipped on a per-token basis, and its pricing sheet has a structure that most launch coverage will flatten into a single headline number. That structure — specifically a context threshold that reprices your entire request, and reasoning tokens that bill at output rates — is where the real money is.

Every figure below was read from OpenAI's own pricing and model documentation on 3 September 2026 and cross-checked against each competing vendor's pricing page. Anything we could not confirm at source is flagged rather than guessed.

How much does GPT-6 Astra cost per million tokens?

Here is the complete rate card, verified 3 September 2026 against OpenAI's API pricing page. All figures are US dollars per million tokens.

TierContextInputCached inputCache writeOutput
Standard≤ 272K$10.00$1.00$12.50$50.00
Standard> 272K$20.00$2.00$25.00$75.00
Batch / Flex≤ 272K$5.00$0.50$6.25$25.00
Batch / Flex> 272K$10.00$1.00$12.50$37.50
Fast mode≤ 272K$20.00$2.00$25.00$100.00
Fast mode> 272K$40.00$4.00$50.00$150.00

Three things in that table deserve more attention than the headline $10/$50.

First, cache writes cost more than fresh input — $12.50 against $10.00. Writing to the cache carries a 25% premium, so a prompt prefix you cache but only reuse once loses money. You need at least two reads for caching to pay for itself.

Second, Batch and Flex are priced identically at 50% of standard. Flex gives you a synchronous request at a slower, best-effort service level; Batch gives you asynchronous processing. Same price, different latency profile — pick on the shape of your workload, not on cost.

Third, Fast mode is a flat 2× multiplier on whatever rate already applies — OpenAI describes it as "priced at 2x the applicable rates." The widely-repeated claim that it delivers 2.5× throughput appears on neither the pricing page nor the model page, so benchmark it yourself before budgeting around it.

For context on the model itself: Astra carries a 1,050,000-token context window with a 922,000-token maximum input, 128,000 max output tokens, and a knowledge cutoff of 30 April 2026.

What is GPT-6 Astra's long-context pricing threshold?

This is the single most decision-relevant number on the page, and it is the one most coverage omits: 272,000 input tokens.

OpenAI's GPT-6 Astra model page states it verbatim: "Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request."

For the full request. That phrase is the whole story. This is not a tiered utility bill where the first 272K tokens bill at the cheap rate and only the overflow bills at the expensive one. Cross the line by a single token and every token in the request — including the 272,000 that were under the threshold — reprices upward. Input doubles, cached input doubles, cache writes double, and output goes up 1.5×.

Work through two nearly identical requests:

Request ARequest B
Input tokens270,000275,000
Output tokens2,0002,000
Input rate$10 / 1M$20 / 1M
Output rate$50 / 1M$75 / 1M
Input cost$2.70$5.50
Output cost$0.10$0.15
Total$2.80$5.65

Request B sends 1.9% more input and costs 102% more. That is a cliff, not a slope, and it means any pipeline whose prompt size varies around a quarter-million tokens will show wildly inconsistent per-request costs for reasons that look inexplicable on a dashboard.

The engineering implication is blunt: instrument a hard token count before you send. A request landing between 260K and 272K is one retrieved document away from doubling in price. Truncate, or split into two sub-272K calls — the arithmetic below shows when that wins.

Do GPT-6 Astra's reasoning tokens bill at output rates?

Yes, and this is the second place naive cost estimates fall apart. OpenAI's reasoning guide is explicit: "While reasoning tokens are not visible via the API, they still occupy space in the model's context window and are billed as output tokens."

So the $50 per million output rate applies to text you never see. Astra supports reasoning effort levels of low, medium, high, xhigh, and max — notably, it does not support the none setting available on other models in the family, so you cannot turn reasoning off entirely.

What that does to a cost estimate: take a request with 10,000 input tokens that returns a 1,500-token visible answer. Estimated naively, that is $0.10 of input plus $0.075 of output — about $0.175. Now suppose the model burned 4,500 reasoning tokens getting there, which is unremarkable at higher effort settings. Real output is 6,000 tokens, costing $0.30, and the request totals $0.40 — roughly 2.3× the naive estimate.

The reasoning-to-answer ratio varies by effort and task, and OpenAI publishes no typical figure, so treat that multiplier as a scenario rather than a benchmark. The actionable point stands regardless: read usage.output_tokens from the API response, not the length of the text you rendered. The gap between them bills at $50 per million.

How does GPT-6 Astra compare with the frontier field on price?

Rates below are standard-tier, per million tokens, verified 3 September 2026 at each vendor's own pricing documentation.

ModelInputOutputCached input
GPT-6 Astra (≤272K)$10.00$50.00$1.00
GPT-6 Astra (>272K)$20.00$75.00$2.00
Claude Fable 5.1$10.00$50.00$0.25
Claude Opus 5$5.00$25.00$0.50
Claude Sonnet 5$2.00$10.00$0.20
GPT-5.6 Sol$4.00$20.00$0.40
GPT-5.6 Terra$2.00$12.00$0.20
GPT-5.6 Luna$0.20$1.20$0.02
Gemini 3.8 Flash$0.75$3.75$0.075
Muse Spark 1.3 (Meta)Open weights — listed at $0.00; you pay for hosting

A few readings. Astra is 2.5× the price of GPT-5.6 Sol, OpenAI's own previous flagship, on both input and output — and 50× Luna's input rate. Against Google, Astra input costs 13× Gemini 3.8 Flash's promotional $0.75, though note that Flash's rate doubles to $1.50/$7.50 on 1 January 2027, so that gap narrows to 6.7× in four months.

Meta's Muse Spark 1.3 is listed by Artificial Analysis at $0.00 in and out because it ships as open weights rather than a metered first-party API. Its Intelligence Index score of 62 is, on that benchmark, marginally above Astra's 61 — a result worth knowing before assuming the price gap tracks capability. Self-hosting has real costs, but they are infrastructure costs, not per-token ones.

One correction to circulating launch coverage: GPT-6 Astra is not currently available on Amazon Bedrock. Bedrock's pricing catalogue lists GPT-5.6 Sol, Terra, Luna, Cyber, GPT-5.5, GPT-5.4, and the gpt-oss models, but Astra does not appear among them as of 3 September 2026. OpenAI's own docs list Chat Completions, Responses, and the Batch endpoint as the API surfaces. If you need Astra through a cloud marketplace today, verify availability directly rather than assuming parity with the 5.6 family.

Why does Claude Fable 5.1 cost the same but bill differently?

Astra and Claude Fable 5.1 are priced identically on the headline: $10 input, $50 output. That is a striking coincidence between two independently-developed frontier models, and it makes the two genuinely comparable on paper.

They are not comparable on cache. Anthropic's model documentation notes that prompt cache reads normally cost 10% of base input price — but 2.5% on Fable 5.1 specifically. That works out to $0.25 per million cached tokens against Astra's $1.00: a 4× difference on the one line item that dominates any workload with a large stable prefix.

If your architecture is a long system prompt plus a big retrieved context reused across many turns — which describes most production RAG and most coding agents — cached reads are the majority of your input volume. At that point the two models stop being equivalently priced. The worked monthly examples below quantify it.

The comparison cuts the other way too: Claude Opus 5 sits at $5/$25, exactly half Astra's rate, while charging $0.50 for cache reads. So Opus 5 is cheaper than Astra on fresh input and output but twice Fable 5.1's cache rate. There is no single winner — the answer depends entirely on your cache hit ratio. We worked through that trade-off in detail in our Fable 5.1 versus Opus 5 comparison.

What does the per-token table miss about real cost?

Per-token rates tell you what a token costs. They do not tell you what a task costs, because models differ enormously in how many tokens they spend reaching an answer — and as established above, Astra's reasoning tokens all bill at $50 per million.

Artificial Analysis publishes a cost-to-run figure that captures this. Running GPT-6 Astra at max effort across their Intelligence Index cost $3,013.30, for a score of 61 (ranked 8th of 202 models evaluated). They also compute a blended rate of $7.70 per million tokens using a 7:2:1 ratio of cached input to fresh input to output — usefully below the $10 headline, because in that mix most input is cached.

The blended figure is a better planning number than the headline rate if you cache heavily. And the $3,013.30 evaluation cost is a reminder that at max effort, a model that thinks expensively can cost several multiples of a cheaper one at identical per-token rates.

What does GPT-6 Astra cost for three realistic monthly workloads?

Arithmetic below uses the verified standard-tier rates. Token volumes are illustrative but sized to plausible production loads.

Workload 1: support assistant, 200,000 requests/month

8,000 input tokens per request (6,000 from a cached prefix, 2,000 fresh), 800 output tokens including reasoning.

Line itemVolumeGPT-6 AstraClaude Fable 5.1Claude Opus 5
Cached input1,200M$1,200$300$600
Fresh input400M$4,000$4,000$2,000
Output160M$8,000$8,000$4,000
Monthly total$13,200$12,300$6,600

The cache differential alone hands Fable 5.1 a $900/month advantage at identical headline pricing. Opus 5 comes in at exactly half of Astra.

Workload 2: long-document analysis, 1,000 documents/month

300,000 input tokens per document — every request is over the 272K line — and 4,000 output tokens.

  • As-is (long-context rates): 300M input × $20 = $6,000, plus 4M output × $75 = $300. Total $6,300.
  • Split into two 150K-token calls: 300M input × $10 = $3,000, plus 8M output × $50 = $400 (output doubles — two calls, two answers). Total $3,400.

Chunking below the threshold saves $2,900 a month, a 46% reduction, even after paying for twice the output. That is the clearest demonstration of why the 272K number matters more than the headline rate for document-heavy pipelines.

Workload 3: overnight classification, 500,000 items/month

1,200 input tokens and 150 output tokens per item, latency-insensitive.

  • Astra, standard: $6,000 input + $3,750 output = $9,750
  • Astra, Batch: $3,000 input + $1,875 output = $4,875
  • GPT-5.6 Luna, standard: $120 input + $90 output = $210

Batch halves the Astra bill. Choosing a right-sized model instead cuts it by 23×. For bulk classification, tier selection matters far less than model selection — a point we expand on in our roundup of the cheapest fast LLM APIs.

How do you reduce a GPT-6 Astra bill?

  1. Stay under 272K input tokens. The highest-leverage single change available. Count tokens before dispatch and truncate, chunk, or route around the threshold. Crossing it costs more than any other pricing decision on this page.
  2. Cache aggressively, but reuse at least twice. Cache writes cost $12.50/M against $10.00 for fresh input, so a single-use cached prefix is a net loss. From the second read onward, at $1.00/M, it is a 90% discount.
  3. Move anything latency-insensitive to Batch or Flex. Both run at exactly 50% of standard. There is no quality difference — only a scheduling one.
  4. Lower reasoning effort where the task allows. Reasoning tokens bill at the full output rate. Dropping from xhigh to medium on routine work cuts invisible output volume directly. Note that none is unavailable on Astra.
  5. Route by task, not by default. Astra at 2.5× GPT-5.6 Sol's price is worth it for genuinely hard reasoning. It is indefensible for classification, extraction, or routing, where Luna at $0.20/$1.20 does the job.
  6. Use Fast mode surgically. A flat 2× on every token, with a throughput benefit OpenAI does not quantify publicly. Reserve it for interactive paths where latency is the product.

What does GPT-6 Astra cost on a ChatGPT plan instead of the API?

OpenAI's model documentation lists GPT-6 Astra as available on the Plus, Pro, Business, and Enterprise ChatGPT plans, with rollout beginning for enterprise customers in the Trusted Access Program. It is not on the free tier.

The economics are structurally different, not merely cheaper. A subscription is a flat fee governed by message caps; the API is metered with no ceiling on usage or spend. For interactive individual work a subscription almost always wins — a single 300K-token analysis costs $5.65 on the API at long-context rates, and a handful of those exceeds a monthly seat. For anything programmatic or embedded in a product, the API is the only option.

We are not quoting subscription dollar figures here: OpenAI's consumer pricing page was not retrievable at the time of writing, and every number in this article comes from a source we opened. Check OpenAI's pricing page for current plan costs.

Which model should you actually pick?

A decision rule, based on the verified numbers above rather than on benchmark scores:

If your prompts sit under 272K tokens and your workload is dominated by a large cached prefix, Claude Fable 5.1 is the cheaper of the two identically-priced frontier models — the 4× cache-read advantage compounds across every request. On mostly-fresh input that advantage evaporates and you are choosing on capability, not cost.

If your prompts routinely exceed 272K tokens, price the chunked alternative before accepting the long-context rate. In our worked example it was 46% cheaper, and the arithmetic favours splitting in most cases where the output is short relative to the input.

And if you are reaching for Astra as a default rather than for a specific hard-reasoning task, you are probably overpaying by 2.5× against GPT-5.6 Sol or 50× against Luna. Frontier pricing is defensible for frontier problems. Most production tokens are not frontier problems.

For a fuller picture of what Astra actually does well, see our complete GPT-6 Astra guide and our Astra versus GPT-5.6 Sol comparison.

FAQ

How much does GPT-6 Astra cost?

GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens on the standard tier, with cached input at $1.00 and cache writes at $12.50. Prompts above 272,000 input tokens are billed at $20 input and $75 output for the entire request. Batch and Flex run at 50% of standard; Fast mode is 2×.

What is GPT-6 Astra's long-context pricing?

Any prompt exceeding 272,000 input tokens is priced at 2× the input and cache rates and 1.5× the output rate — $20 input, $2 cached input, $75 output per million. Critically, the repricing applies to the full request, not just the tokens above the threshold, so crossing the line by one token roughly doubles the cost of the whole call.

Is GPT-6 Astra more expensive than GPT-5.6 Sol?

Yes, by 2.5× on both dimensions. GPT-5.6 Sol costs $4 per million input and $20 per million output, against Astra's $10 and $50. Cached input is $0.40 on Sol against $1.00 on Astra. Sol is also available on Amazon Bedrock, which Astra currently is not.

Is GPT-6 Astra cheaper than Claude Fable 5.1?

Not on balance. Both list at $10 input and $50 output, so headline rates are identical. But Fable 5.1 reads cached tokens at $0.25 per million against Astra's $1.00 — a 4× difference. On any workload with a substantial cached prefix, Fable 5.1 works out cheaper; on all-fresh input the two are equivalent.

Does GPT-6 Astra have batch pricing?

Yes. The Batch API runs at 50% of standard rates: $5 input, $0.50 cached input, $6.25 cache write, and $25 output per million at short context. The Flex tier is priced identically to Batch but serves requests synchronously at a best-effort service level, so you can pick on latency rather than cost.

Is GPT-6 Astra included in ChatGPT Plus?

Yes. OpenAI lists Astra as available on the Plus, Pro, Business, and Enterprise ChatGPT plans, with enterprise rollout starting through the Trusted Access Program. It is not on the free tier. Subscription access is governed by message caps rather than per-token billing, so it is generally cheaper than the API for interactive individual use.

Do GPT-6 Astra's reasoning tokens cost extra?

They bill at the full output rate of $50 per million, even though they are never returned to you. OpenAI's documentation states reasoning tokens "are billed as output tokens." Astra supports effort levels from low through max but does not support disabling reasoning entirely, so always budget from the API's reported output token count rather than the visible answer length.