Claude Sonnet 5.5 vs GPT-6.1 Sol (2026): Benchmarks and Real Cost

Quick answer. Claude Sonnet 5.5 (Sep 28, 2026) and GPT-6.1 Sol (Sep 29, 2026) both list at $2 input / $10 output per 1M tokens. Sonnet 5.5 scores 56 vs 52 on the Artificial Analysis Intelligence Index and wins most of its agentic and knowledge-work evals, but at max effort it costs $7.60 per task vs $0.72 for Sol. Pick Sonnet for quality, Sol for cost per task.

Anthropic shipped Claude Sonnet 5.5 on September 28, 2026. OpenAI answered the next day at DevDay with GPT-6.1 Sol, a point release that replaced GPT-6 Sol a week after it launched. Both are the "mid-tier" models of their families, both are now in GitHub Copilot, and both cost exactly the same per token: $2 per million input tokens and $10 per million output tokens.

Identical sticker prices make the usual comparison useless. What actually separates these two models is how they score on the same tests, how many tokens they burn to finish a task, what cached input costs, and what happens to your bill past 272K tokens of context. This article lines up the numbers that are comparable (mostly from Artificial Analysis, which ran both models on the same harness), flags the ones that are not, and works through a real cost example before giving a verdict.

What are Claude Sonnet 5.5 and GPT-6.1 Sol?

Claude Sonnet 5.5 (claude-sonnet-5-5) is Anthropic's newest Sonnet, sitting below Claude Opus 5.5 in the lineup. Anthropic's docs describe it as "the best combination of speed and intelligence", with adaptive thinking, five effort levels (low to max) and a default effort of high. Anthropic claims it is "30%+ faster" than Sonnet 5 and costs "up to 30% less" per task at the same per-token price. It is on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. For the full breakdown, see our Claude Sonnet 5.5 complete guide.

GPT-6.1 Sol (gpt-6.1-sol) is OpenAI's mid-tier model under GPT-6 Astra. OpenAI's model page says it is built for "complex coding, computer use, and professional work". It supports reasoning effort from low through max (default medium), has an April 30, 2026 knowledge cutoff, and runs on the Responses API, Batch, and Chat Completions (without tool calling on Chat Completions). Artificial Analysis says it lands 1 point below GPT-6 Astra on its Intelligence Index at less than a quarter of Astra's cost per task. Our GPT-6.1 Sol complete guide covers it in depth, and the GPT-6 Sol and Luna guide covers the model it replaced.

Sonnet 5.5 vs GPT-6.1 Sol: specs and pricing side by side

SpecClaude Sonnet 5.5GPT-6.1 Sol
Release dateSep 28, 2026Sep 29, 2026 (DevDay)
API model IDclaude-sonnet-5-5gpt-6.1-sol
Input price (per 1M)$2.00$2.00
Output price (per 1M)$10.00$10.00
Cached input / cache read (per 1M)$0.20$0.10
Cache write (per 1M)$2.50 (5-min), $4.00 (1-hour)$2.50
Context window1M tokens1,050,000 tokens (922K max input)
Long-context surchargeNone listed; full 1M billed at standard rate>272K input tokens: 2x input and cache rates, 1.5x output, for the whole request
Max output128K (300K on Batch API, beta header)128K
Batch discount50% off input and output50% off (Batch and Flex: $1 / $5)
Effort levelslow, medium, high (default), xhigh, maxlow, medium (default), high, xhigh, max
Knowledge cutoffJun 2026Apr 30, 2026
GitHub Copilot plansPro, Pro+, Max, Business, EnterprisePro+, Max, Business, Enterprise

Sources: Anthropic model docs and announcement, OpenAI model page and pricing page, GitHub Changelog. OpenAI's pricing page also lists a "Fast" tier at $4 / $20.

The three pricing differences that matter

  • Cache reads: Sol's cached input is half the price of Sonnet's ($0.10 vs $0.20). For agent loops that resend a big system prompt and repo context every turn, this adds up.
  • Long context: Once a Sol request crosses 272K input tokens, the entire request is repriced at 2x input (and 2x cache) and 1.5x output. Sonnet 5.5 bills its full 1M window at the standard rate. Above 272K, Sol's cache reads become $0.20, the same as Sonnet's.
  • Copilot access: Sonnet 5.5 is on the Copilot Pro plan; GPT-6.1 Sol starts at Pro+.

Which scores higher on benchmarks?

Here is the catch with vendor benchmark tables: Anthropic and OpenAI mostly report different benchmarks, at different effort levels, sometimes on different versions of the same test. Anthropic reports OSWorld 2.1; OpenAI reports OSWorld 2.0. Anthropic reports Terminal-Bench 4.0; OpenAI does not report a Terminal-Bench 4.0 number for Sol at all. So the only clean head-to-head comes from Artificial Analysis (AA), which ran both models through the same evaluations.

Independent head-to-head (Artificial Analysis, max effort)

Benchmark (AA harness)Claude Sonnet 5.5 (max)GPT-6.1 Sol (max)Winner
AA Intelligence Index5652Sonnet +4
Terminal-Bench 4.0 (agentic terminal coding)64%56%Sonnet +8
GDPval-AA v2.1 (professional knowledge work, Elo)18441575Sonnet +269
AA-Briefcase v1.1 (Elo)18111564Sonnet +247
AutomationBench-AA71%65%Sonnet +6
Humanity's Last Exam55%53%Sonnet +2
SciCode61%54%Sonnet +7
GDP.pdf (document reasoning)26%31%Sol +5
CritPt (physics reasoning)31%32%Sol +1
AA-Omniscience Index (knowledge reliability, -100 to 100)3242Sol +10
AA-LCR v1.1 (long-context reasoning)83%83%Tie
Cost per Intelligence Index task$7.60$0.72Sol, ~10.5x cheaper
Output speed139.1 tok/s66.2 tok/sSonnet
Time to first token438.42 s272.81 sSol

Source: Artificial Analysis model comparison page, both models at max effort. AA's own launch article calls Sonnet 5.5 "#2" on the Index; its model page currently lists it at #3 of 222, so the rank has shifted as new models were added. The score of 56 is consistent across both.

The pattern is clear. On agentic coding, business automation and knowledge-work evals, Sonnet 5.5 is ahead, sometimes by a lot (the GDPval-AA gap is over 250 Elo). Sol wins three of the ten Index evals: AA-Omniscience (a knowledge-reliability score that rewards correct answers and penalizes hallucinations), GDP.pdf, and CritPt by a single point. Long-context reasoning (AA-LCR) is a tie. The long time-to-first-token figures reflect how much both models think at max effort; they drop sharply at lower effort (see below).

Vendor-reported numbers (not directly comparable)

BenchmarkClaude Sonnet 5.5GPT-6.1 SolReported by
Terminal-Bench 4.070.6% (Anthropic's max-effort config)not reportedAnthropic
OSWorld (computer use)80.1% on OSWorld 2.1 (partial credit)71.4% on OSWorld 2.0 (max effort)Anthropic / OpenAI via Vellum
CursorBench 4.055.5%not reportedAnthropic
FrontierCode 1.1 (Main)46.2% (max)not reportedAnthropic
DeepSWE v1.171.0% (reported by The New Stack, via Vellum)75.2% (high)OpenAI / third party
AutomationBench 1.0.644.7% ($1.14/task)36.0% ($0.30/task)Compiled by Vellum
GDP.pdfnot reported32.0% (~$0.38/task)OpenAI via Vellum

Treat this table as context, not a scoreboard. Note the Terminal-Bench 4.0 gap between Anthropic's own 70.6% and AA's independent 64% for the same model: vendor harnesses and effort settings move numbers by several points. The OSWorld rows are different test versions. DeepSWE is the one place a third party puts Sol ahead, and that Sonnet figure is secondhand, so we would not lean on it until someone runs both on the same harness.

Which is cheaper per task? The token-usage problem

This is the most important section of the comparison. Same price per token does not mean same price per job, because the two models spend very different amounts of tokens to finish the same task.

Artificial Analysis measured Sonnet 5.5 at max effort using about 193K output tokens per Intelligence Index task, which it called "the highest token use we have measured", roughly 60% above Opus 5.5 (max) and about 7x GPT-6 Astra (max). Across the whole index, Sonnet 5.5 (max) generated 410M output tokens; GPT-6.1 Sol (max) generated 67M. That is about 6x the output for the same set of tasks.

Cost per task at matched effort levels (Artificial Analysis)

Model and effortAA Intelligence IndexCost per taskOutput tokens (whole index)SpeedTTFT
Sonnet 5.5 (max)56$7.60410M139.1 tok/s438.42 s
Sonnet 5.5 (xhigh)52$2.74100M103.1 tok/s33.50 s
Sonnet 5.5 (high, API default)47$1.0850M93.5 tok/s14.47 s
GPT-6.1 Sol (max)52$0.7267M66.2 tok/s272.81 s
GPT-6.1 Sol (xhigh)51$0.3936M63.9 tok/s108.86 s
GPT-6.1 Sol (high)50$0.3225M64.6 tok/s57.97 s

Source: Artificial Analysis model pages for each variant. AA also lists GPT-6.1 Sol (low) at $0.13 per task.

Read the table by intelligence level, not by effort label. At a score of 52, Sonnet 5.5 needs xhigh effort and costs $2.74 per task; GPT-6.1 Sol reaches 52 at max effort for $0.72. That is roughly 3.8x cheaper for Sol at equal Index score. At Sonnet's API default (high), it scores 47 for $1.08, while Sol at high scores 50 for $0.32. On AA's aggregate measure, Sol gives you more intelligence per dollar at every point below Sonnet's max-effort ceiling.

The flip side: Sonnet 5.5 is the only one of the two that reaches 56. If your task needs that last 4 points, Sol has no higher setting to buy it with, and Sonnet at max still costs more per task than Opus 5.5 at max ($5.98, per AA's Opus 5.5 model page). For that tier, compare against Opus directly in our Claude Opus 5.5 vs GPT-6 Sol comparison.

Sonnet also streams faster once it starts (93 to 139 tok/s vs roughly 64 to 66 tok/s for Sol), and its time to first token at high effort (14.47 s) is far shorter than Sol's at high (57.97 s). For interactive IDE work, that latency gap is felt more than the cost gap.

Worked example: what does a month of agent runs cost?

Here is a simple model for a team running a coding agent. We use AA's measured cost per task, which already reflects each model's token appetite, then layer on the two list-price effects AA's figure does not isolate: caching and long context.

Scenario A: 20,000 agent tasks a month, AA cost per task

SetupAA IndexCost per task20,000 tasks / month
Sonnet 5.5, high (default)47$1.08$21,600
Sonnet 5.5, xhigh52$2.74$54,800
Sonnet 5.5, max56$7.60$152,000
GPT-6.1 Sol, high50$0.32$6,400
GPT-6.1 Sol, max52$0.72$14,400

Simple multiplication of AA's per-task figures. AA's tasks are an evaluation mix, not your workload, so treat these as relative, not absolute. The ratios are the point: matching Sol's max-effort Index score with Sonnet costs about 3.8x more.

Scenario B: one long-context agent turn at list prices

Now take a single agent turn with a 150K-token cached prefix (system prompt plus repo context), 20K fresh input tokens, and 15K output tokens. Total input is 170K, under Sol's 272K threshold.

  • Sonnet 5.5: 150K cached x $0.20 = $0.030; 20K input x $2 = $0.040; 15K output x $10 = $0.150. Total $0.220.
  • GPT-6.1 Sol: 150K cached x $0.10 = $0.015; 20K input x $2 = $0.040; 15K output x $10 = $0.150. Total $0.205.

At equal output, Sol's cheaper cache saves about 7% here. Now push the same turn to 400K input (380K cached, 20K fresh), which crosses Sol's 272K line:

  • Sonnet 5.5: 380K x $0.20 = $0.076; 20K x $2 = $0.040; 15K x $10 = $0.150. Total $0.266.
  • GPT-6.1 Sol (2x input and cache, 1.5x output): 380K x $0.20 = $0.076; 20K x $4 = $0.080; 15K x $15 = $0.225. Total $0.381, about 43% more than Sonnet.

So Sol's per-token advantage flips once you live above 272K tokens of context. But these list-price examples hold output constant, and output is where the real gap is. If Sonnet produces several times more output tokens for the same job (as AA measured at max effort), that swamps both the cache discount and the long-context surcharge. The honest rule: measure tokens per task on your own workload at the effort level you will ship. Both APIs return usage in every response.

How to call each model (and set effort)

Effort is the main cost lever on both models. Anthropic's docs recommend starting Sonnet 5.5 at medium for well-specified agentic coding, high for harder work, and using xhigh or max only where your evals show a gain.

import anthropic

client = anthropic.Anthropic()

response = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=128000,
    messages=[{"role": "user", "content": "Refactor this function and add tests: ..."}],
    output_config={"effort": "medium"},
)

for block in response.content:
    if block.type == "text":
        print(block.text)
print(response.usage)

GPT-6.1 Sol uses the reasoning.effort field on the Responses API:

from openai import OpenAI

client = OpenAI()

response = client.responses.create(
    model="gpt-6.1-sol",
    reasoning={"effort": "high"},
    input=[{"role": "user", "content": "Refactor this function and add tests: ..."}],
)

print(response.output_text)
print(response.usage)

Two Sonnet 5.5 gotchas from Anthropic's model docs: setting temperature, top_p or top_k to a non-default value returns a 400 error, and forced tool use returns an error. If you are porting a Sonnet 5 integration, read the migration guide before swapping the model string.

Which is better for coding in September 2026?

On the independent coding signal we have, Sonnet 5.5 is ahead: 64% vs 56% on AA's Terminal-Bench 4.0 run, plus Anthropic's own 55.5% on CursorBench 4.0 and 46.2% on FrontierCode 1.1, neither of which OpenAI reports for Sol. GitHub's changelog positions Sonnet 5.5 for "well-scoped everyday work like building features and fixing bugs" and says it matched Sonnet 5 on coding "while using significantly fewer steps, tokens, and tool calls" in their testing. GitHub describes GPT-6.1 Sol as strong for "agentic coding and terminal workflows" with "efficient token use".

OpenAI's strongest coding claim is DeepSWE v1.1, where Sol at high effort (75.2%) roughly matches GPT-6 Astra at about a fifth of the cost. That is a strong result, but it is not a head-to-head against Sonnet on a shared harness.

Our read: if you are picking "the best coding model of September 2026" by quality alone at this price tier, Sonnet 5.5 edges it. If you are picking by quality per dollar for high-volume agent runs, GPT-6.1 Sol does. For how these models fit into Claude Code, Codex, Cursor and Copilot workflows, see our coding agent comparison.

Verdict: pick Sonnet 5.5 or GPT-6.1 Sol?

Pick Claude Sonnet 5.5 if...

  • You need the highest score at the $2 / $10 tier. It wins seven of AA's ten Index evals, ties one, and only Sonnet reaches an Index of 56.
  • Your work is knowledge-work heavy (documents, analysis, business automation). The GDPval-AA and AA-Briefcase gaps are about 250-270 Elo.
  • You regularly send more than 272K tokens of context. Sonnet bills its whole 1M window at standard rates.
  • You want lower latency in an IDE. At high effort, Sonnet's first token arrives in about 14 s vs about 58 s for Sol, and it streams faster.
  • You are on GitHub Copilot Pro, where Sol is not available.

Pick GPT-6.1 Sol if...

  • Cost per task matters more than the last few points. At an equal AA Index score of 52, Sol costs $0.72 per task vs $2.74 for Sonnet.
  • You run high-volume agents with large cached prompts under 272K tokens. Sol's $0.10 cache read is half Sonnet's.
  • Knowledge reliability without tools matters. Sol scores 42 vs 32 on AA's Omniscience Index, and it also leads on GDP.pdf document reasoning (31% vs 26%).
  • You already build on the OpenAI Responses API or Batch/Flex tiers ($1 / $5).

Or use both

A common pattern: route routine, well-scoped tasks to GPT-6.1 Sol at high effort and escalate the hard ones to Sonnet 5.5 at xhigh or max. Since both are in Copilot and both have simple effort controls, this is a model-picker change, not a rewrite. Whatever you choose, run an effort sweep on your own tasks: Anthropic says Sonnet 5.5's effort levels are recalibrated from Sonnet 5, and Artificial Analysis found GPT-6.1 Sol uses 10-30% more output tokens than GPT-6 Sol at each effort setting.

FAQ

Is Claude Sonnet 5.5 better than GPT-6.1 Sol?

On Artificial Analysis's head-to-head at max effort, yes: Sonnet 5.5 scores 56 vs 52 on the Intelligence Index and wins Terminal-Bench 4.0, GDPval-AA, AA-Briefcase, AutomationBench-AA, SciCode and Humanity's Last Exam. GPT-6.1 Sol wins AA-Omniscience, GDP.pdf and CritPt, and wins on cost per task by a wide margin.

Do Sonnet 5.5 and GPT-6.1 Sol cost the same?

Per token, almost. Both are $2 input and $10 output per 1M tokens. Sol's cached input is $0.10 vs Sonnet's $0.20, and Sol charges 2x input and 1.5x output for requests over 272K input tokens. Per task, they differ a lot: AA measured $7.60 for Sonnet (max) vs $0.72 for Sol (max).

Why is Sonnet 5.5 so expensive per task at max effort?

It thinks at great length. Artificial Analysis measured about 193K output tokens per Index task at max effort, the highest it has recorded. Dropping to high or medium effort cuts this sharply: at high, AA measured $1.08 per task.

What is the context window of each model?

Claude Sonnet 5.5 has a 1M-token context window with 128K max output (300K on the Batch API with a beta header). GPT-6.1 Sol has a 1,050,000-token window with up to 922K input and 128K output.

Are both available in GitHub Copilot?

Yes. Sonnet 5.5 arrived on September 28, 2026 for Copilot Pro, Pro+, Max, Business and Enterprise. GPT-6.1 Sol arrived on September 29, 2026 for Pro+, Max, Business and Enterprise. Both appear in the model picker in VS Code, Visual Studio, JetBrains, Xcode, Eclipse, Copilot CLI and the coding agent.

Which is the best coding model in September 2026?

At the $2 / $10 tier, Sonnet 5.5 has the stronger independent coding result (64% vs 56% on AA's Terminal-Bench 4.0). GPT-6.1 Sol is the better value for high-volume agent work. Above this tier, Opus 5.5 and GPT-6 Astra remain the heavier options.

Can I trust the vendor benchmark tables?

Use them carefully. Anthropic reported 70.6% on Terminal-Bench 4.0 for Sonnet 5.5, while AA measured 64% on its own harness. The two vendors also use different benchmark versions (OSWorld 2.1 vs 2.0). Compare like with like, which in practice means independent evaluations such as Artificial Analysis.

Sources

If your team is building on these models, Codersera can help you hire vetted remote developers.