GPT-6 Luna is the model you reach for when you need to run something a million times a day and the per-call cost is the whole design constraint. It shipped alongside GPT-6 Sol on 22 September 2026, and at $0.10 per 1M input tokens it is currently the cheapest input rate of any frontier-lab reasoning model with a million-token context.
This is the dedicated Luna deep-dive: the published specification, where the price sits inside the GPT-6 family, whether the jump from GPT-5.6 Luna is worth making, and — the part most coverage skips — what Luna is genuinely bad at. Every number below comes from OpenAI's model documentation, the API changelog, or Artificial Analysis's independent testing.
What is GPT-6 Luna?
GPT-6 Luna is a reasoning model that accepts text and image input and returns text. OpenAI describes it in its model index as "our most efficient model for focused, high-volume tasks" — the efficiency tier below GPT-6 Sol (coding and agentic work) and GPT-6 Astra (the frontier tier).
Here is the full published specification:
| Property | GPT-6 Luna |
|---|---|
| Model string | gpt-6-luna |
| Released | 22 September 2026 |
| Context window | 1,050,000 tokens |
| Max input tokens | 922,000 |
| Max output tokens | 128,000 |
| Knowledge cutoff | 18 May 2026 |
| Input modalities | Text, image |
| Output modalities | Text |
| Reasoning effort levels | none, low, medium (default), high, xhigh, max |
| Endpoints | Chat Completions, Responses, Batch |
| Not supported | Realtime, Live, Assistants, fine-tuning, embeddings |
| Input price | $0.10 / 1M tokens |
| Output price | $0.50 / 1M tokens |
Two details in that table are worth slowing down on, because almost every secondary write-up gets them wrong.
The context window is not the input limit. Luna's context window is 1,050,000 tokens, but the maximum input is 922,000 and the maximum output is 128,000. Those two numbers add up to exactly 1,050,000 — the window is a shared budget, and OpenAI reserves the full output allowance out of it. If you plan to ask for a long generation, you cannot also fill 1M tokens of input. Size your prompts against 922,000, not 1,050,000.
reasoning.effort: none is supported. This matters more than it sounds. GPT-6 Astra, the top-tier model, does not accept none — its documentation lists only low through max. Luna does, which means you can run it as a plain fast non-reasoning model with no thinking-token overhead at all. For classification and extraction work, that is the setting you want, and it is the single biggest lever on both cost and latency.
Luna supports streaming, structured outputs, function calling, prompt caching and image input, plus the full hosted-tool set: web search, file search, code interpreter, hosted shell, apply patch, skills, computer use, MCP and tool search.
What does GPT-6 Luna cost, and how does it compare inside the GPT-6 family?
Luna's full rate card, per 1M tokens:
| Rate | Price / 1M tokens |
|---|---|
| Input | $0.10 |
| Cached input | $0.01 |
| Cache write | $0.125 |
| Output | $0.50 |
| Batch / Flex input | $0.05 |
| Batch / Flex output | $0.25 |
Batch and Flex processing run at 50% of standard pricing. Fast mode costs 2x. Cached input at $0.01 per 1M is a 10x discount on reads — on a workload with a stable system prompt, caching is doing more for your bill than any model choice.
Now the family spread, which is the genuinely useful framing:
| Model | Input / 1M | Output / 1M | Cached input | Knowledge cutoff | Supports none? |
|---|---|---|---|---|---|
| GPT-6 Luna | $0.10 | $0.50 | $0.01 | 18 May 2026 | Yes |
| GPT-6 Sol | $2.00 | $10.00 | $0.20 | 20 Apr 2026 | Yes |
| GPT-6.1 Sol | $2.00 | $10.00 | $0.10 | — | — |
| GPT-6 Astra | $10.00 | $50.00 | $1.00 | 30 Apr 2026 | No |
Luna is 20x cheaper than Sol on both input and output, and 100x cheaper than Astra on input. All four share the same 1,050,000-token context window and the same 922,000 / 128,000 input-output split, so you are not buying context when you move up the family — you are buying reasoning quality only.
Luna also has the freshest knowledge in the family: 18 May 2026, roughly a month ahead of Sol's 20 April and Astra's 30 April. If your task depends on recent facts rather than hard reasoning, that is a real point in Luna's favour.
The 272K long-context cliff
Luna has the same pricing threshold as the rest of the GPT-6 line, and it will bite you if you do not plan for it. Prompts with more than 272,000 input tokens are billed at 2x input and cache rates, and 1.5x output, for the full request — not just the tokens above the line.
So a 280,000-token prompt costs $0.20 per 1M input, not $0.10, and output jumps to $0.75 per 1M. If you are near the boundary, trim below 272K — retrieval that keeps prompts under the threshold pays for itself immediately at this scale.
For the broader family context — how Sol and Luna were positioned at launch and where each one fits — see our combined GPT-6 Sol and Luna guide, and the GPT-6 Astra deep-dive for the frontier tier.
How does GPT-6 Luna compare to GPT-5.6 Luna?
This is the question with live search demand and, until now, no page answering it directly. GPT-5.6 Luna launched 9 July 2026; GPT-6 Luna replaced it ten weeks later.
| GPT-6 Luna | GPT-5.6 Luna | Change | |
|---|---|---|---|
| Released | 22 Sep 2026 | 9 Jul 2026 | +10 weeks |
| Input / 1M | $0.10 | $0.20 | −50% |
| Output / 1M | $0.50 | $1.20 | −58% |
| Cached input / 1M | $0.01 | $0.02 | −50% |
| Cache write / 1M | $0.125 | $0.25 | −50% |
| Context window | 1,050,000 | 1,050,000 | Unchanged |
| Max input / output | 922,000 / 128,000 | 922,000 / 128,000 | Unchanged |
| Knowledge cutoff | 18 May 2026 | 16 Feb 2026 | +3 months |
| Reasoning levels | none–max | none–max | Unchanged |
| 272K cliff | 2x in / 1.5x out | 2x in / 1.5x out | Unchanged |
And the independent numbers, from Artificial Analysis — all three measured at max reasoning effort, on Intelligence Index v4.3.2 (October 2026). AA rebased the index from v4.1.1 to v4.3.2 and there is no conversion factor between the two, so only compare these against other v4.3.2 scores:
| Model | AA Intelligence Index v4.3.2 | Blended price / 1M | Output speed | Time to first token |
|---|---|---|---|---|
| GPT-6 Luna (Max) | 38.12 | $0.08 | 140.7 tok/s | 96.0 s |
| GPT-5.6 Luna (Max) | 37.32 | $0.17 | 126.1 tok/s | 106.4 s |
| GPT-6 Sol (Max) | 47.63 | $1.54 | 107.1 tok/s | 107.8 s |
Read that honestly: GPT-6 Luna is under one index point smarter than GPT-5.6 Luna. 38.12 versus 37.32 on Intelligence Index v4.3.2 is inside the noise of a composite benchmark. It is 12% faster on output and shaves ten seconds off time-to-first-token, which is a real but modest improvement.
The upgrade is not a capability jump. It is a price jump — the same intelligence for half the money, and blended cost drops from $0.17 to $0.08 per 1M on Artificial Analysis's 7:2:1 cache-hit/input/output mix. AA's own cost-to-run figure says the same thing from the other direction: $0.07 per Index task for GPT-6 Luna (Max) against $0.18 for GPT-5.6 Luna (Max) and $1.04 for GPT-6 Sol (Max) — $122.05 to run the entire v4.3.2 Index on Luna. That is the entire story of this release, and it is a good story if you are running Luna at volume.
Should you migrate? Yes, and it is close to a no-brainer. Every spec that could break your integration is identical — same context window, same input and output ceilings, same reasoning-effort options, same endpoints, same tools, same 272K threshold. Change the model string from gpt-5.6-luna to gpt-6-luna, re-run your evals to confirm nothing regressed on your specific task, and your bill halves. There is no migration tax here beyond the eval run. For the previous generation's full line-up, including Terra which has no GPT-6 equivalent, see our GPT-5.6 Sol, Terra and Luna guide.
What is GPT-6 Luna actually good for?
A model at $0.10 / $0.50 with an Intelligence Index of 38.12 at max effort (v4.3.2) has a clear shape, and pretending otherwise wastes your money. Here is where Luna earns its place.
Classification and routing. Set reasoning.effort: none, use structured outputs to pin the response to an enum, and you have a near-free classifier that costs about $0.10 to label a million short inputs. This is the single best use of Luna and the reason it exists.
Extraction over large documents. With 922,000 usable input tokens you can push an entire contract set, log dump or codebase into one call and pull structured fields out. At $0.10 per 1M input, reading a 250,000-token document costs two and a half cents — and staying under 272K keeps you off the 2x cliff.
High-volume summarisation. Support tickets, transcripts, changelogs, review feeds. Output is the expensive side at $0.50 per 1M, so cap max_output_tokens deliberately; summarisation is naturally output-light, which is why the economics work.
Cheap tool-calling and lightweight agent steps. Luna supports function calling and the full hosted-tool set, so it works as the cheap worker inside an agent loop — the step that decides which tool to call, while a Sol or Astra call handles the step that actually needs judgement.
Anything behind a stable system prompt. Cached input at $0.01 per 1M is the real unlock. A long fixed instruction block read ten thousand times costs almost nothing.
Where GPT-6 Luna will disappoint you
Being specific here is more useful than hedging:
- Serious coding. Luna (Max) scores 38.12 on Intelligence Index v4.3.2 against GPT-6 Sol (Max)'s 47.63. Nine and a half points on a composite that includes coding evaluations is not a rounding error. OpenAI positions Sol, not Luna, as the model "built to power complex coding and agentic workflows." Use Luna to triage a bug report; use Sol to fix the bug.
- Long multi-step agentic runs. Small per-step quality gaps compound. A loop that is 90% reliable per step is 35% reliable over ten steps. Cheap steps that fail are not cheap.
- Latency-critical interactive UX. The 96-second time-to-first-token Artificial Analysis measured is at
maxreasoning effort, and you would drop tononeorlowfor interactive use — but do not assume "the cheap model" means "the snappy model" by default. Reasoning effort, not tier, controls latency. - Hard reasoning, research, and novel problem-solving. This is what the 100x-more-expensive Astra tier is for. Luna at
maxeffort burns a lot of thinking tokens to get to 38.12; paying for reasoning on a model that is not built to reason is the most common way teams waste money on the cheap tier.
How does GPT-6 Luna compare to other cheap-tier models?
Luna's competition is the flash tier from every other lab. Current list prices per 1M tokens:
| Model | Input / 1M | Output / 1M | Context | AA Index v4.3.2 |
|---|---|---|---|---|
| GPT-6 Luna | $0.10 | $0.50 | 1,050,000 | 38.12 (Max) |
| GLM-5.3-Flash | $0.15 | $0.50 | 1,048,576 | 41.81 |
| Qwen3.8 Omni-Flash | $0.15 | $0.47 | 1,000,000 | Not scored by AA |
| DeepSeek V4.1 Flash | $0.30 | $1.20 | 1,048,576 | 39.46 (Max) |
| Gemini 3.8 Flash | $0.75 | $3.75 | 1,048,576 | 40.93 (High) |
Luna has the lowest input price in the group. On output it is effectively tied with GLM-5.3-Flash and Qwen3.8 Omni-Flash, and it is 7.5x cheaper than Gemini 3.8 Flash on output — which is a reminder that "Flash" means very different things at different labs.
It also has the lowest Intelligence Index of every scored model in that table. GLM-5.3-Flash is 41.81 on v4.3.2 against Luna (Max)’s 38.12, for five cents more per 1M input and an identical output rate — 3.7 index points ahead at effectively the same price. Gemini 3.8 Flash (High) is 40.93 and DeepSeek V4.1 Flash (Max) is 39.46. Qwen3.8 Omni-Flash has no AA record at all, so treat its quality as unmeasured rather than assuming it tracks its Flash sibling. Luna wins the cheap-tier price race and comes last in the cheap-tier quality race; if your workload is quality-sensitive at the margin, GLM-5.3-Flash is the better default and the price gap is rounding error.
Price is not the only axis, though. Qwen3.8 Omni-Flash handles more input modalities than Luna's text-and-image. The open-weight options — GLM-5.3-Flash, DeepSeek V4.1 Flash, the Qwen line — can be self-hosted, which changes the calculus entirely at sustained high volume and for data-residency requirements. Luna's advantages are the 10x cache-read discount, the full OpenAI hosted-tool surface, and reasoning.effort: none as a first-class setting. For a broader cross-vendor view, see our comparison of the cheapest fast LLM APIs.
What's the difference between GPT-6 Luna and GPT-6 Luna Pro?
Less than the name suggests, and this trips people up.
There is no separate GPT-6 Luna Pro model card in OpenAI's documentation — the URL simply does not exist. "Luna Pro" is a serving configuration, not a different model. It is the same underlying GPT-6 Luna weights served with reasoning.mode set to pro, which produces higher-quality answers on complex tasks by thinking harder.
The per-token price is identical: $0.10 input, $0.50 output, batch at $0.05 / $0.25. But pro mode spends substantially more reasoning tokens per request, and reasoning tokens bill as output. So Luna Pro costs the same per token and considerably more per answer, with no fixed multiplier to plan against — it depends on how much the model decides to think.
The practical implication: if you are already calling Luna at reasoning.effort: high or max and want more quality, Luna Pro is the next step. If you are running Luna at none or low because the economics are the point, Luna Pro is the opposite of what you want — and at that quality level you should be comparing against GPT-6.1 Sol instead, which is purpose-built for harder work.
How do you use GPT-6 Luna?
The model string is gpt-6-luna. It is available on Chat Completions, Responses and Batch. Here is a working Responses API call with reasoning disabled — the configuration you want for classification:
curl https://api.openai.com/v1/responses \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-luna",
"input": "Classify this ticket as billing, bug, or feature_request: My card was charged twice this month.",
"reasoning": { "effort": "none" },
"max_output_tokens": 16
}'
And the same thing in Python, with structured outputs to guarantee a parseable label:
from openai import OpenAI
client = OpenAI()
resp = client.responses.create(
model="gpt-6-luna",
input="Classify this ticket: My card was charged twice this month.",
reasoning={"effort": "none"},
max_output_tokens=16,
text={
"format": {
"type": "json_schema",
"name": "ticket_label",
"schema": {
"type": "object",
"properties": {
"label": {
"type": "string",
"enum": ["billing", "bug", "feature_request"],
}
},
"required": ["label"],
"additionalProperties": False,
},
}
},
)
print(resp.output_text)
Three things to set on every production Luna call:
- Pick reasoning effort explicitly. The default is
medium, which means you are paying for thinking tokens you may not need.noneis free of that overhead entirely. - Cap
max_output_tokens. Output is 5x the price of input. An unbounded ceiling on a model that can emit 128,000 tokens is an unbounded bill. - Use Batch for anything not interactive. Half price, same model, and the tier-5 batch queue limits on this family run into the billions of tokens.
Rate limits scale by account tier:
| Tier | Requests / min | Tokens / min |
|---|---|---|
| 1 | 500 | 500,000 |
| 2 | 5,000 | 2,000,000 |
| 3 | 5,000 | 4,000,000 |
| 4 | 10,000 | 10,000,000 |
| 5 | 30,000 | 180,000,000 |
Tier 5 at 180M tokens per minute is a notably high ceiling — higher than Sol's 40M — which reflects what OpenAI expects Luna to be used for.
Should you use GPT-6 Luna?
A simple decision rule.
If you are currently on GPT-5.6 Luna, move. The specs are identical, the quality is a hair better, and the price is half. The only work is re-running your evals.
If you are choosing a cheap tier fresh, Luna has the lowest input price available and the best cache-read discount in its class. Take it unless you need modalities it does not have, or you need open weights you can host yourself.
If you are reaching for Luna to save money on work that is actually hard — multi-step agents, real coding, novel reasoning — stop. The 9.5-point Intelligence Index gap (v4.3.2, both at max) to Sol is the gap that matters, and a cheap wrong answer costs more than an expensive right one. Use Luna for the volume and Sol or Astra for the judgement; the family is priced the way it is precisely so you can split the work.
If you are building production systems on top of models like these and want engineers who have already worked through the cost and reliability trade-offs, Codersera can help you extend your team.
FAQ
What is GPT-6 Luna?
GPT-6 Luna is OpenAI's efficiency-tier reasoning model in the GPT-6 family, released on 22 September 2026. It accepts text and image input, returns text, has a 1,050,000-token context window, and is built for high-volume, latency-sensitive tasks like classification, extraction and summarisation. Its model string is gpt-6-luna.
How much does GPT-6 Luna cost?
$0.10 per 1M input tokens and $0.50 per 1M output tokens. Cached input is $0.01 per 1M and cache writes are $0.125 per 1M. Batch and Flex processing run at half price: $0.05 input and $0.25 output. Prompts above 272,000 input tokens are billed at 2x input and 1.5x output rates.
Is GPT-6 Luna better than GPT-5.6 Luna?
Marginally on quality, dramatically on price. Artificial Analysis scores GPT-6 Luna (Max) at 38.12 on Intelligence Index v4.3.2 versus 37.32 for GPT-5.6 Luna (Max) — within benchmark noise — but it costs half as much ($0.10/$0.50 against $0.20/$1.20) and runs about 12% faster. Every other spec is identical, so migration is low-risk.
What is GPT-6 Luna's context window?
1,050,000 tokens total, split into a maximum of 922,000 input tokens and 128,000 output tokens. Those add to exactly 1,050,000, so the window is a shared budget — you cannot fill 1M tokens of input and still request a long generation. Prompts above 272,000 input tokens trigger 2x input pricing.
What's the difference between GPT-6 Luna, Sol and Astra?
They are the same context window at three price and capability points. Luna is $0.10/$0.50 for high-volume work, Sol is $2/$10 for coding and agentic workflows, and Astra is $10/$50 for the hardest reasoning. Luna is 20x cheaper than Sol and 100x cheaper than Astra on input. Astra is also the only one that does not support reasoning.effort: none.
Is GPT-6 Luna good for coding?
Not for serious coding. Luna (Max) scores 38.12 on Artificial Analysis's Intelligence Index v4.3.2 against GPT-6 Sol (Max)'s 47.63, and OpenAI explicitly positions Sol — not Luna — as the model built for complex coding and agentic workflows. Luna is fine for triaging bug reports, labelling diffs or summarising changelogs, but use Sol to actually write or fix code.
What is GPT-6 Luna Pro?
GPT-6 Luna Pro is not a separate model — it is the same GPT-6 Luna weights served with reasoning.mode set to pro, which produces better answers on complex tasks by spending far more reasoning tokens. The per-token price is identical at $0.10/$0.50, but because reasoning tokens bill as output, the cost per answer is meaningfully higher.
Does GPT-6 Luna support image input?
Yes. GPT-6 Luna accepts text and image input and returns text only. There is no audio or video input, and no image generation output — though the hosted image_generation tool is available for tool-based use. Supported endpoints are Chat Completions, Responses and Batch; Realtime, Live, Assistants, fine-tuning and embeddings are not supported.