Meta listed Muse Spark 1.3 on 2 September 2026, four weeks after 1.2. It is the first Muse Spark release to land within two points of a frontier Anthropic model on a neutral harness — and it does so at a quarter of the price, with video input that Claude Opus 5 does not accept at all.
The boring version of this comparison — "cheap model loses to expensive model" — stopped being true at 1.3. But the gap that remains sits in exactly the place buyers care about most. Every figure below was pulled on 3 September 2026 from the source named beside it, with vendor-reported and independently-run numbers kept apart, because they are not the same kind of evidence.
What do the two models actually give you?
Specifications from the live OpenRouter model API and Anthropic's own model documentation, both read on 3 September 2026.
| Muse Spark 1.3 | Claude Opus 5 | |
|---|---|---|
| Listed | 2 Sep 2026 | 24 Jul 2026 |
| Context window | 1,048,576 | 1,000,000 |
| Max output tokens | 943,718 | 128,000 |
| Input / 1M | $1.25 | $5.00 |
| Output / 1M | $4.25 | $25.00 |
| Cached input / 1M | $0.15 | $0.50 |
| Input modalities | text, image, file, video, audio* | text, image, file |
| Built-in web search | $2.50 / 1k calls | $10 / 1k calls |
| Second SKU | Contributor: $0.10 / $0.20 | Batch: $2.50 / $12.50 |
| Knowledge cutoff | Not published | May 2026 |
*Meta documents audio on Muse Spark 1.3 as "not fully supported" — see the multimodal section below.
Three of those rows are structural rather than incremental: Muse Spark 1.3 is 4× cheaper on input and 5.9× cheaper on output, it takes video natively, and its output ceiling is 7.4× higher. Everything else is close enough to be a wash — the context windows differ by 4.9%, which will never decide a purchase.
Anthropic's documentation adds a detail worth carrying into the coding section: Opus 5 is now described as the model "for complex agentic coding and enterprise work," while Claude Fable 5.1 is pointed at "demanding reasoning and long-horizon agentic work." Opus 5 is no longer the top of Anthropic's own range.
What do independent benchmarks say?
Artificial Analysis puts Muse Spark 1.3 at an Intelligence Index of 61 against Claude Opus 5 at 63. On the model's own page the max-effort variant scores 62 and ranks 6th of 636 models tracked. The context that matters: Muse Spark 1.2 scored 57 on the same index. A four-point generational gain in four weeks has cut the gap to Opus 5 from six points to two.
| Artificial Analysis | Muse Spark 1.3 | Muse Spark 1.2 | Claude Opus 5 |
|---|---|---|---|
| Intelligence Index | 61 | 57 | 63 |
| Output speed | 186 tok/s | 172 tok/s | 56 tok/s |
| Time to first token | 34.2 s | 17.9 s | 81.7 s |
| Cost to run the index | $0.55 | $0.40 | $2.34 |
Terminal-Bench 4.0 — the verified agentic-coding leaderboard — tells the opposite story, and it is the single most important table in this article:
| Rank | Model | Harness | Resolution rate | Run cost |
|---|---|---|---|---|
| 1 | Claude Opus 5 (max) | Claude Code | 51.8% ± 3.4 | $5,969 |
| 2 | Claude Fable 5 | Claude Code | 44.5% ± 3.8 | $7,265 |
| 3 | GLM-5.3 | Claude Code | 41.8% ± 3.2 | $2,728 |
| 4 | GPT-5.6 Sol | Codex | 37.3% ± 3.8 | $2,542 |
| 5 | Claude Opus 4.8 | Claude Code | 23.6% ± 3.6 | $6,481 |
| 6 | GPT-5.6 Terra | Codex | 21.5% ± 3.3 | $1,734 |
| 7 | Grok 4.6 | Grok Build | 20.3% ± 3.1 | $3,592 |
No Muse Spark model appears on that board at any rank. Not 1.1, not 1.2, not 1.3. The same holds on SWE-bench Verified, where the official submission archive contains no entry for either Muse Spark or Opus 5 — so on the two hardest verified coding harnesses, Opus 5 has one first-place result and Muse Spark has nothing to check.
On human preference the ranking inverts again. The LMArena text leaderboard places Muse Spark 1.2 (xHigh) 5th at 1499±10 from 3,240 votes, ahead of Claude Opus 5 High in 9th at 1493±5 from 35,174 votes. Muse Spark 1.3 has not accumulated enough votes to be listed yet. Read that gap carefully: a ±10 confidence interval against a ±5 one, on a tenth of the sample, is a tie dressed up as a rank.
Vals AI, on its own common harness, had Opus 5 at 67.21% against Muse Spark 1.2 at 57.05% as of 1 September; 1.3 was not yet scored.
The honest summary: on composite intelligence and human preference the two are within noise of each other. On verified, reproducible agentic coding, only one of them has a number at all.
Is Muse Spark 1.3 cheaper than Opus 5 in practice?
Yes, but by less than the price list implies. Three worked examples at list price. The Contributor column uses the meta/muse-spark-1.3-contributor SKU at $0.10 input / $0.20 output.
| Workload | Muse Spark 1.3 | Contributor | Claude Opus 5 |
|---|---|---|---|
| Document job: 1M in, 100k out | $1.68 | $0.12 | $7.50 |
| Generation job: 50k in, 500k out | $2.19 | $0.11 | $12.75 |
| Agent loop: 10M cached in, 200k out | $2.35 | $0.06 | $10.00 |
List price says 4.5× to 5.8× cheaper, and 62× to 167× cheaper on the Contributor tier. The measured figure is lower. Artificial Analysis records $0.55 to run its full index on Muse Spark 1.3 against $2.34 on Opus 5 — a real-world gap of 4.3×, not 5.9×.
The discrepancy has a cause you can see in the same dataset: Muse Spark 1.3 emitted 120 million output tokens across the index run against a 72 million median, which Artificial Analysis flags as "very verbose." A model that is 5.9× cheaper per output token but produces 1.7× more of them does not deliver 5.9× savings. Budget for roughly 4×, not 6×.
The Contributor tier is a different proposition entirely. At $0.10/$0.20 it is 125× cheaper on output than Opus 5, which changes what is economically sensible to automate rather than just trimming a bill. The trade is documented rather than implied — Meta's pricing page describes it as discounted pricing "in exchange for permission to use your prompts and completions to train future Meta models," against the Standard tier where prompts and completions "are not used to train Meta models."
There is a second limit that matters more than most people expect. Contributor is capped at 100 requests per minute against Standard's 3,000, and Meta applies those limits per team rather than per API key — so you cannot shard your way around it. That makes Contributor a prototyping and batch-experimentation tier, not a serving tier: the 125× discount is real, but you can only push about a thirtieth of the traffic through it. Most of the headline saving is not usable in production. For the data terms in full, see our Muse Spark 1.3 complete guide and, for the same tier on Meta's coding agent, what Muse Code's Contributor tier actually costs you. Do not put client code through it without reading those.
What can Muse Spark 1.3 see that Opus 5 cannot?
Video and images. Audio is listed but comes with a caveat Meta itself publishes, and it is the part of this comparison most likely to be reported wrongly.
Start with what is unambiguous. Anthropic's documentation states that all current Claude models support text and image input — there is no video path on Opus 5 at any price. For anything arriving as moving footage, Muse Spark 1.3 is the only one of the two that can take it, and both Muse SKUs carry the identical modality set, so the cheap tier is not stripped down here.
Audio is where the OpenRouter listing and Meta's own docs disagree. The listing advertises audio input on 1.3. Meta's developer documentation says, verbatim:
"Audio understanding in Muse Spark 1.3 is currently not fully supported, and response quality for requests including audio content may be degraded."
Meta points audio work at Muse Spark 1.2 or its dedicated muse-voice-transcribe-1.0 model instead, and every audio example in its own guides is written against muse-spark-1.2 even though 1.3 is the default elsewhere. Meta's video-understanding page applies the same caveat to embedded soundtracks, so narrated video and screen recordings with commentary are caught by it too — the frames are fine, the audio track is not.
The advantage is real, then, but narrower than the spec sheet implies:
| Input | Muse Spark 1.3 | Claude Opus 5 |
|---|---|---|
| Images, PDFs | Yes | Yes |
| Silent video / screen capture | Yes | No |
| Narrated video | Frames yes, audio degraded | No |
| Standalone audio | Degraded — use 1.2 or Voice Transcribe | No |
So a silent screen recording of a bug, a UI walkthrough or a product demo goes to Muse Spark 1.3 in one call, and cannot go to Opus 5 at all. A support call or an hour of standup audio should go to Muse Spark 1.2 or Voice Transcribe — not 1.3, and not Opus 5, which needs a separate speech-to-text vendor first.
It is an unusual regression to ship knowingly, and it reads as a signal: 1.3 is a coding and agentic point release, and audio was not what it was tuned for.
Does a 943,718-token output limit actually matter?
Less than the headline suggests, and the arithmetic has a catch. Muse Spark 1.3's ceiling is 943,718 output tokens against 128,000 for Opus 5, a 7.4× difference. But that output has to fit inside the same 1,048,576-token context window as your input. Ask for the maximum output and you have 104,858 tokens of input budget left — about a tenth of the window. It is a split of one pool, not extra capacity bolted on.
Two caveats keep this honest. Anthropic's Message Batches API raises the Opus 5 ceiling to 300,000 tokens with the output-300k-2026-03-24 beta header, so for asynchronous work the real gap is 3.1×, not 7.4×. And a single generation approaching a million tokens is rarely the right shape for a task — coherence degrades long before the limit.
Where the headroom earns its place is in avoiding truncation logic. Whole-repository migrations, bulk structured extraction, exhaustive translation passes and large synthetic-data runs all cross 128,000 output tokens routinely. On Opus 5 those need chunking and continuation handling — code you write, test and debug. On Muse Spark 1.3 they are one call. An engineering-time saving, not a capability unlock.
Which is better for coding and agentic work?
Claude Opus 5, and it is not close. Opus 5 holds first place on Terminal-Bench 4.0 at 51.8% ± 3.4 across 330 trials — a run that consumed 6.5 billion tokens and $5,969, with an average trial duration of about 80 minutes. That is a long-horizon agentic benchmark run to completion by a third party and published with a confidence interval. Muse Spark 1.3 has no entry on that board, none on SWE-bench Verified, and Meta has published no launch benchmark post for 1.3 at all.
That last point is a change from 1.2. In August, Meta published a benchmark chart with 1.2 in second place behind Opus 5, plus a methodology note conceding the comparison was "not harness-identical to the leaderboard" — we took it apart in our Muse Spark 1.2 benchmarks vs Claude Opus 5 analysis. For 1.3 there is no equivalent: Meta's AI blog index still shows "Introducing Muse Spark 1.1" from 9 July as its most recent Muse Spark post, and the developer model page carries pricing and three positioning claims — "trained for long-horizon, agentic workflows," "competitive coding performance," "native multimodal perception" — with no numbers behind them.
So there is no vendor claim to verify. Artificial Analysis's index of 61 is the only independent capability signal on 1.3 today, and a composite index is not an agentic-coding result. If you need a model to drive a terminal agent unsupervised for an hour, Opus 5 is the one with proof.
Where is each model genuinely worse?
Muse Spark 1.3 is worse at:
- Verified agentic coding. Absent from Terminal-Bench 4.0 and SWE-bench Verified. Zero reproducible evidence, four weeks after the previous version also had none.
- Latency. 34.2 s to first token on Artificial Analysis — nearly double Muse Spark 1.2's 17.9 s. Whatever 1.3 gained in intelligence, it spends in thinking time before it answers.
- Verbosity. 120M output tokens per index run against a 72M median. Costs more than the price list implies and produces more text to review.
- Audio. Listed as an input type, but Meta documents it as "not fully supported" on 1.3 and routes audio work back to 1.2 — a regression against the previous version.
- Track record. Our earlier analysis of the 1.1 launch found Meta's claimed 80.0% on Terminal-Bench 2.1 sitting above the independently verified 76.2% ± 1.2 — a small gap, but in a consistent direction across releases.
Claude Opus 5 is worse at:
- Price. 4.3× more expensive per measured task, before you consider the Contributor tier.
- Throughput. 56 tokens/second against 186 — a 3.3× difference that users feel directly in interactive tools.
- Latency at max effort. 81.7 s to first token, the slowest of the three configurations compared here.
- Media inputs. No audio, no video, at any tier.
- Output ceiling. 128,000 tokens synchronously, 300,000 on batch.
- Position in its own range. Anthropic's docs now route "demanding reasoning and long-horizon agentic work" to Fable 5.1, not Opus 5.
Which should you pick?
One question decides it: does anything unsupervised depend on the output being right?
If yes — an agent editing a repository, a production migration, a code review nobody reads afterwards — use Claude Opus 5. It is the only one of the two with a verified long-horizon agentic result, and the price gap is trivial next to an unsupervised mistake in a codebase. If no — classification, summarisation, extraction, drafting, anything with a human or a test suite in between — use Muse Spark 1.3. Two index points is not worth 4.3× the bill on work that gets checked.
| Use case | Pick | Why |
|---|---|---|
| Unsupervised terminal / repo agent | Claude Opus 5 | Only verified Terminal-Bench 4.0 result of the pair |
| Production code review | Claude Opus 5 | Verified coding evidence; error cost dominates price |
| Video or screen-capture input | Muse Spark 1.3 | Opus 5 has no video path at any price |
| Audio or call recordings | Neither — use Muse Spark 1.2 | Meta flags audio on 1.3 as degraded |
| High-volume batch classification | Muse Spark 1.3 | 4.3× cheaper measured, 3.3× faster output |
| Single generation over 300k tokens | Muse Spark 1.3 | 943,718 ceiling vs 300,000 on Opus 5 batch |
| Interactive, latency-sensitive UX | Muse Spark 1.3 | 186 tok/s vs 56, despite slower first token |
| Non-sensitive prototyping | Contributor tier | 125× cheaper output, but capped at 100 RPM per team |
| Client or regulated data | Claude Opus 5 | No data-rights trade in the pricing |
A practical hybrid: route the bulk stage to Muse Spark 1.3 and the verification stage to Opus 5. On the agent-loop example above that is roughly $2.35 plus a small Opus 5 check rather than $10.00 end to end, and the verified model still sits where the decision happens.
Recheck this in four weeks. Muse Spark shipped three versions since mid-July and gained four index points in the last four weeks alone. The two-point gap is the narrowest it has been, and it is closing from one direction.
FAQ
Is Muse Spark 1.3 better than Claude Opus 5?
Not overall, but it is close. Artificial Analysis scores Muse Spark 1.3 at 61 on its Intelligence Index against 63 for Claude Opus 5. On verified agentic coding the gap is much wider: Opus 5 ranks first on Terminal-Bench 4.0 at 51.8%, while no Muse Spark model appears on that leaderboard at all. Muse Spark 1.3 wins on price, speed and media inputs.
Is Muse Spark 1.3 cheaper than Opus 5?
Substantially. List price is $1.25/$4.25 per million input/output tokens against $5.00/$25.00 for Opus 5 — 4× and 5.9× cheaper respectively. Measured end to end the gap narrows to about 4.3×, because Muse Spark 1.3 is verbose and emits roughly 1.7× the median number of output tokens per task. The Contributor tier is cheaper again at $0.10/$0.20.
Which is better for coding?
Claude Opus 5, clearly, on current evidence. It holds first place on Terminal-Bench 4.0 at 51.8% ± 3.4 across 330 trials, run independently with the Claude Code harness. Muse Spark 1.3 has no entry on Terminal-Bench or SWE-bench Verified, and Meta published no benchmark chart with the 1.3 release, so there is no vendor claim to check either.
Does Muse Spark 1.3 support audio?
Only nominally. Audio is a listed input type, but Meta's own docs say "audio understanding in Muse Spark 1.3 is currently not fully supported, and response quality for requests including audio content may be degraded," and point you to Muse Spark 1.2 or muse-voice-transcribe-1.0 instead. The caveat also covers soundtracks embedded in video. Video and image input, by contrast, work fully — and Claude Opus 5 has no video path at all.
What is the Muse Spark Contributor tier?
A second SKU at $0.10 input and $0.20 output per million tokens — about 125× cheaper on output than Opus 5 — with identical context, output limit and modalities. Meta grants the discount "in exchange for permission to use your prompts and completions to train future Meta models." It is also rate-limited to 100 requests per minute against Standard's 3,000, applied per team, which makes it a prototyping tier rather than a serving tier.
Should I switch from Opus 5 to Muse Spark 1.3?
Switch the workloads where output gets checked by a human or a test suite — bulk classification, summarisation, extraction, drafting. Keep Opus 5 for anything unsupervised that touches a codebase, since it is the only one of the two with a verified long-horizon agentic result. A hybrid split, bulk on Muse Spark and verification on Opus 5, captures most of the saving.
How big is the context and output difference?
Context windows are nearly identical: 1,048,576 tokens for Muse Spark 1.3 against 1,000,000 for Opus 5. Max output differs sharply — 943,718 versus 128,000 synchronously, though Anthropic's Message Batches API raises Opus 5 to 300,000 with a beta header. Note that Muse Spark's output shares the same window, so a maximum-length generation leaves only about 104,858 tokens for input.
Where does Muse Spark 1.3 rank against other models?
Artificial Analysis places the max-effort variant sixth of 636 models tracked, at an index of 62. On Terminal-Bench 4.0 there is no Muse Spark entry at all. For the wider Muse family and how the versions differ, see our Muse Spark complete guide.
If you are still on the previous generation, our Muse Spark 1.2 benchmarks comparison covers the upgrade case, and the Claude Opus 5 launch guide has the full Anthropic-side detail.
Figures verified on 3 September 2026 against the OpenRouter model API, Meta's developer model, pricing and video-understanding docs, Anthropic's model documentation, Artificial Analysis, the Terminal-Bench 4.0 leaderboard, the SWE-bench Verified submission archive, LMArena and Vals AI. Benchmark leaderboards move; recheck before committing a migration.