Quick answer. Meta's own launch benchmarks put Muse Spark 1.2 second to Claude Opus 5 on all three coding evaluations it published — 82.9% vs 86.7% on Terminal-Bench 2.1, 59.3% vs 65.0% on DeepSWE, and 70.6% vs 79.4% on Meta's internal benchmark. It is still absent from every verified agentic-coding leaderboard, and LiveBench scores its agentic coding at 57.6 — a regression from 1.1.
Update — Muse Spark 1.3 is out. Meta shipped Muse Spark 1.3 on 2 September 2026. This page stays focused on the 1.2 benchmark claims and whether they held up; for the current version, read Muse Spark 1.3 vs Claude Opus 5.
When a vendor launches a model, the benchmark chart is a marketing asset. Reading it well means asking three questions: who ran it, who was in the comparison set, and does anyone independent get the same number?
For Muse Spark 1.2, launched alongside Muse Code on August 5, 2026, the answers are unusually interesting — because Meta published a chart on which it loses.
What did Meta actually claim?
Three coding benchmarks, all run by Meta, all showing Muse Spark 1.2 in second place behind Claude Opus 5.
| Model | Terminal-Bench 2.1 | DeepSWE v1.1 | Meta Internal Coding Bench |
|---|---|---|---|
| Claude Opus 5 | 86.7% | 65.0% | 79.4% |
| Muse Spark 1.2 | 82.9% | 59.3% | 70.6% |
| GPT-5.6 Terra | 81.8% | 64.8% | 65.4% |
| Grok 4.5 | 81.6% | 56.6% | — |
| Muse Spark 1.1 | 76.2% | 53.0% | 68.3% |
One row that chart cannot have: Grok 4.6, which xAI shipped on August 12, a week after Meta published. There is no Meta-run figure for it, and xAI's own launch numbers are not droppable into this column — xAI reports Terminal-Bench v3.0 (26%), a different and much harder version than the 2.1 scores above. Independent 2.1 runs of Grok 4.6 did exist as of August 2026 (Artificial Analysis 88.39%, Vals 78.28%), but they come from different harnesses again, so they belong beside this table, not inside it. Both figures are also now frozen: Artificial Analysis retired Terminal-Bench 2.1 from its Intelligence Index when it rebased the Index to v4.3.2 in September 2026, replacing it with Terminal-Bench 4.0, so there is no current refresh of that 88.39%.
A fourth chart, covering kernel-generation work, places Muse Spark 1.2 fourth of six — so across every benchmark Meta chose to publish, it leads none of them. Meta's methodology note also concedes the comparison is "not harness-identical to the leaderboard" and "may not reflect these models' best performance."
Publishing a chart you lose is a defensible choice — it signals confidence that the price-to-performance ratio is the real argument. But it does mean the headline is settled: Muse Spark 1.2 is not the strongest coding model available. Meta says so.
The generational improvement is real, though. Against Muse Spark 1.1, that is +6.7 points on Terminal-Bench and +6.3 on DeepSWE in roughly four weeks. That pace is the thing competitors should be watching, not the absolute position.
Who was left out of the comparison?
This is where the chart gets less generous.
Meta benchmarked against GPT-5.6 Terra — OpenAI's mid-tier model — rather than Sol, the top of the range. Sol scores meaningfully higher on Terminal-Bench than any model in Meta's table. Swapping in the flagship would have pushed Muse Spark 1.2 to third.
Claude Fable 5 is also absent from the coding charts, and the Meta Internal Coding Bench row omits several competitors entirely.
None of this is unusual — every vendor picks its comparison set — but it is the difference between "second-best coding model" and "second-best among the models Meta chose to include."
Does independent testing agree?
No, and the gap is large enough to matter.
The core problem with vendor benchmarks is the harness. A benchmark score measures a model plus the agent scaffold wrapped around it. Vendors run their own model inside their own optimised harness, which inflates results relative to a neutral setup.
Vals AI runs every model through one common harness. As of its August 12 update, Muse Spark 1.2 sits 5th on the Vals Index at 71.88% — behind Claude Fable 5, Claude Opus 5, Kimi K3 and GPT-5.6 Sol — while carrying the lowest cost per test in the top five, about $0.69. That is a strong showing for a budget model, and it climbed sharply in the week after launch.
The same board moved again on August 12, when Grok 4.6 entered 6th at 71.82% — 0.05 points behind Muse Spark 1.2, which is a tie in everything but sort order. Artificial Analysis reads the pair the other way round, and still does on the current basket: on Intelligence Index v4.3.2 (October 2026) it puts Grok 4.6 at 42.84, or 44.31 at high reasoning effort, against Muse Spark 1.2's 39.58. When two neutral harnesses disagree about which of two models is ahead, the honest reading is that neither one is.
But the composite flatters it. On LiveBench it scores 78.0 overall with strong reasoning (90.0) and maths (91.2) — and just 57.6 on agentic coding, its weakest column against peers and a regression from Muse Spark 1.1's 58.5. Agentic coding is the one capability Muse Code is sold on.
Artificial Analysis places Muse Spark 1.2 at 39.58 on Intelligence Index v4.3.2 (captured 5 October 2026) against 50.78 for Claude Opus 5 at max reasoning effort — an 11.2-point gap. That is wider than it looked on the old v4.1.1 basket, where the same pair read 57 against 63. Note that the two numbers are not convertible: AA rebased the Index in September 2026, swapping the component evals (it now runs AA-Briefcase, long-context reasoning, AutomationBench, CritPt, GDP-PDF, GDPval, Humanity's Last Exam, Omniscience, SciCode and Terminal-Bench 4.0), so every score on the board moved without any model changing. We have dropped the rank positions we previously quoted — AA's field has grown past 200 measured models since, and a rank against an old denominator is misleading even when the score behind it is right.
Most importantly, Muse Spark 1.2 still does not appear on any verified agentic-coding leaderboard — not Terminal-Bench, SWE-bench, SWE-rebench or Aider. Its gains are all on composite and preference boards.
And the comparison cuts both ways, which is worth saying plainly: Claude Opus 5 is also absent from the verified Terminal-Bench board. The top verified Anthropic entry is Opus 4.8 at 78.9%. So Meta’s headline "82.9 versus 86.7" is vendor claim against vendor claim, not vendor against verified. Vals also discloses that both Opus 5 and Fable 5 used Claude Opus 4.8 as a refusal fallback in its Terminal-Bench runs; counting those nine affected passes as failures drops Opus 5 from 84.64% to 81.27%.
Is there a track record to judge this against?
Yes, and it is the most useful single data point in this whole exercise.
At the Muse Spark 1.1 launch in July 2026, Meta claimed 80.0% on Terminal-Bench 2.1. Independent verification subsequently returned 76.2% ± 1.2.
The claimed figure sat above the upper bound of the verified confidence interval. Not catastrophically — but it is a systematic direction, not noise. And note that Meta's own launch table for 1.2 lists Muse Spark 1.1 at 76.2%, the verified number, not the 80.0% it originally claimed. The generational gain is being measured against a corrected baseline while the new number remains unverified.
Applying the same correction informally to 1.2 would land it near 79%, which is roughly where Artificial Analysis independently put it on Terminal-Bench 2.1 in August 2026. Two independent methods converging on the same discount is worth noting — though AA has since dropped Terminal-Bench 2.1 for 4.0, so that convergence cannot be re-checked on current data.
To Meta's credit, its published methodology concedes the harness "may not reflect these models' best performance" — a caveat most vendors omit.
How verifiable are Meta's numbers?
Less than you would want. Meta published its benchmark results as chart images rather than as tabular data, and the accompanying methodology document is largely a table of contents. There is no released harness configuration, no per-task breakdown and no raw results file — so the figures cannot be independently recomputed, only re-run from scratch by a third party using a different scaffold.
The clearest illustration of why that matters: Muse Spark 1.1 now has three different Terminal-Bench 2.1 scores in circulation — 76.2%, 78% and 80.0% — depending entirely on who ran it and in which harness. Same model, same benchmark, a four-point spread. When a single model produces that much variance, a 3.8-point gap between two different models is not a reliable ranking signal.
Where does Muse Spark 1.2 genuinely win?
At launch, cost efficiency — and it was not close. Two months on, that is the claim that has aged worst, and it is worth being blunt about.
In August 2026 the case was solid. On evaluations that measure cost alongside capability, Muse Spark 1.2 was the cheapest model in the top tier — roughly $0.69–0.70 per test on Vals AI's common harness, against competitors several times that.
On Artificial Analysis's Intelligence Index v4.3.2 (October 2026) that edge has gone. Running the Index costs $0.97 per task on Muse Spark 1.2 for a score of 39.58. Three models now beat it on both axes simultaneously: GLM 5.3 Flash scores 41.81 at $0.25 per task, MiMo-V2.6-Pro scores 46.32 at $0.13, and GPT-6.1 Sol scores 51.83 at $0.72 — a 12-point capability lead for about a quarter less money. Muse Spark 1.2's price never changed; the cheap tier caught up and went past it.
The structural argument survives even though the ranking does not. A budget model carrying a small capability deficit is still the right buy for bounded, verifiable, high-volume work. In August 2026 that model was Muse Spark 1.2. In October 2026 it is not — which is the case for re-checking the cost-per-task column before every model decision rather than inheriting last quarter's answer.
It also posted the launch cycle's largest jump on GDPval-AA v2, a general knowledge-work evaluation, climbing about 260 Elo points to roughly 1631 and 5th place.
Throughput is genuinely strong, and has held up better than the price argument. Launch-window measurement put it around 165 tokens per second, with real-world tests reporting averages near 191; Artificial Analysis now measures Muse Spark 1.2 at a median 270.1 output tokens per second with an 11.7-second time to first token, which is fast for its class and roughly twice the throughput of Muse Spark 1.3 at max effort (145.9 tok/s, 47.9 s). Speed is the one axis on which 1.2 still beats its own successor.
What did not improve: raw knowledge went slightly backwards, with small declines on SciCode and Humanity's Last Exam. This is a coding-specialised point release, and the specialisation shows in both directions.
What does this look like in practice?
The most credible independent hands-on test so far captures both sides precisely.
On the upside: Muse Code audited 222 pull requests in under five minutes for about ten cents, against roughly $32 for the same job on a frontier model. That is a genuine order-of-magnitude shift in what is economically sensible to automate.
On the downside: in the same session, the agent spent three minutes researching a Google project that does not exist, then built its entire integration plan on that fabrication and could not recover. The tester's verdict — that you cannot trust it for longer-running work — is the single most useful sentence written about the release.
Those two results are consistent with each other, and with the benchmarks. Bounded, verifiable, high-volume tasks are where a cheap, fast, slightly-weaker model wins outright. Long-horizon autonomous work, where a single early error compounds across hours, is where the capability gap gets expensive.
How should you read the numbers?
- Treat vendor benchmarks as an upper bound. Vendor-run scores measure the model plus a harness tuned for it. Expect a few points of regression on a neutral setup.
- Check the comparison set before the bars. Terra rather than Sol changes the ranking without changing a single number.
- Weight verified over claimed. Where independent harnesses exist, they are the number to plan around.
- Benchmark deltas under 5 points rarely survive contact with your codebase. Your language, repo size, test coverage and prompts move results more than the gap between second and fourth place.
- Measure cost per solved task, not per token. A weaker model burns more turns. Once you account for retries, a 12–21x token discount compresses to low single digits on real work.
For the wider leaderboard picture, see our AI agent benchmark roundup.
The verdict
Muse Spark 1.2 is a strong second-tier coding model priced like a budget one. It is not the frontier, by Meta's own accounting, and independent harnesses put it further from the frontier than Meta's chart suggests.
That was a perfectly good product in August 2026, and the rate of improvement — a meaningful jump in about four weeks — was more strategically significant than the position. Two things have changed since. Meta shipped 1.3, which on Artificial Analysis's current basket scores 48.09 against 1.2's 39.58. And the budget tier moved: on Intelligence Index v4.3.2 several cheaper models now score higher than 1.2, so the price-to-performance argument that was its whole case no longer holds on its own numbers.
If you are choosing on capability alone for long-horizon agentic work, Meta's own launch charts still point at Claude Opus 5 — and AA's current data widens rather than narrows that gap, 50.78 against 39.58. If you are choosing on cost per solved task, neither model is the answer in October 2026; check the current board.
For what that means in daily use, see Muse Code vs Claude Code.
FAQ
Is Muse Spark 1.2 better than Claude Opus 5?
No. On all three coding benchmarks Meta published at launch, Claude Opus 5 scored higher — 86.7% vs 82.9% on Terminal-Bench 2.1, 65.0% vs 59.3% on DeepSWE, and 79.4% vs 70.6% on Meta's internal benchmark.
What is Muse Spark 1.2's Terminal-Bench score?
Meta claims 82.9% on Terminal-Bench 2.1, but that figure is vendor-run and Muse Spark 1.2 does not appear on the official verified Terminal-Bench leaderboard. On Vals AI's common harness it now sits 5th on the Vals Index at 71.88%, with the lowest cost per test in the top five. LiveBench scores its agentic coding at 57.6.
Why do independent benchmarks differ from Meta's?
Benchmark scores measure a model plus its agent harness. Vendors optimise their own harness, which inflates results relative to a neutral common scaffold like the one Vals AI uses.
Did Meta overstate Muse Spark 1.1's benchmarks?
Meta claimed 80.0% on Terminal-Bench 2.1 for Muse Spark 1.1; independent verification returned 76.2% ± 1.2, above the confidence interval's upper bound. Meta's 1.2 launch chart now lists 1.1 at the verified 76.2%.
What is Muse Spark 1.2's context window?
1,048,576 tokens (1M), with maximum output around 943,718 tokens. Meta has not published the architecture, parameter count or knowledge cutoff.
Is Muse Spark 1.2 good value?
It was at launch; it no longer is. In August 2026 it was the cheapest model in the top tier at roughly $0.69–0.70 per test on Vals AI. On Artificial Analysis's Intelligence Index v4.3.2 (October 2026) it costs $0.97 per task for a score of 39.58, and GLM 5.3 Flash (41.81 at $0.25), MiMo-V2.6-Pro (46.32 at $0.13) and GPT-6.1 Sol (51.83 at $0.72) all beat it on capability and cost at once. Its remaining edge is speed: 270.1 tok/s median output.
Can I download Muse Spark 1.2 weights?
No. It shipped closed-weights, but on August 10, 2026 Zuckerberg announced Meta will open-source Muse Spark 1.2's weights. They are not out yet and no date has been given. Meta did release Muse Glimmer, a 30B Apache-2.0 open-weights model, on the same day.
Should I switch to Muse Spark 1.2 for coding?
Not in October 2026. The pattern is right — a cheap, fast, slightly-weaker model is the correct buy for bounded, high-volume, verifiable work — but Muse Spark 1.2 is no longer the cheapest model in that band, and Muse Spark 1.3 supersedes it on capability at the same list price. For long-horizon autonomous work, the capability gap and documented hallucination failures still argue for a frontier model.