Quick answer. Muse Spark is Meta's closed-weights multimodal reasoning family from Meta Superintelligence Labs. Four versions have shipped: 1.0 (April 2026, consumer-only), 1.1 (9 July), 1.2 (5 August) and 1.3 (2 September 2026). All API versions cost $1.25 input and $4.25 output per million tokens, with a $0.10/$0.20 Contributor tier that trains on your data.
Muse Spark moves faster than almost any other model line in the market. Four releases in five months, two of them four weeks apart, plus a separate open-weights sibling, a terminal coding agent, an image model and a speech model — all under the same "Muse" umbrella, all shipped since April 2026.
This page is the family reference: what Muse Spark is, every version and what changed, what each one costs today, how the closed-weights question actually stands, and where the rest of the Muse line fits. Version-specific deep dives are linked throughout. Every price, context window and date below was re-checked against OpenRouter's live model API and Meta's published model pages on 3 September 2026. Every Artificial Analysis figure was re-captured on 5 October 2026 and is stated against Intelligence Index v4.3.2 — AA rebased the Index from v4.1.1 in September 2026, swapping the component evals, so older AA scores for these models are not comparable and cannot be converted.
What is Muse Spark?
Muse Spark is a natively multimodal reasoning model from Meta Superintelligence Labs (MSL) — the division Meta stood up in 2025 under Alexandr Wang, the former Scale AI CEO who is now Meta's Chief AI Officer, following Meta's multibillion-dollar investment in Scale AI. "Muse" is the family name; "Spark" is the frontier reasoning line within it.
The strategically important part was never the architecture. It was that Muse Spark is Meta's first proprietary model with no downloadable weights. Meta built the Llama brand on open weights anyone could run; Muse Spark is a hosted service you reach through Meta's apps and a paid API. VentureBeat's framing at the April launch — "Goodbye, Llama" — is still the cleanest summary of what changed.
Three concrete properties define how it behaves in practice, all confirmed on the live model listings:
- Reasoning is mandatory. Every Muse Spark SKU forces reasoning on — you cannot turn it off, only dial it between
minimal,low,medium(the default),highandxhigh. This has a direct billing consequence covered below. - Input is genuinely multimodal. Text, image, video, file (including PDFs) and audio in; text out. Very few frontier models accept all five. The audio leg is the weak one on 1.3 — Meta flags it as not fully supported, and both Artificial Analysis and OpenRouter now list 1.3 without any audio or speech input modality at all, while AA still lists 1.2 with speech.
- It is a moderated endpoint. Every Muse Spark listing is flagged moderated, unlike the open-weights Muse Glimmer, which is not.
Which Muse Spark versions exist, and when did each ship?
Four numbered releases, plus two Contributor SKUs. Muse Spark 1.0 never got a public API — it was a consumer feature inside Meta AI, which is why it has no listing on any model marketplace.
| Version | Shipped | Context | Max output | API price (in / out per 1M) | What it was for |
|---|---|---|---|---|---|
| Muse Spark 1.0 | April 2026 | ~262K | — | No public API | Meta's first closed model; consumer-only inside Meta AI |
| Muse Spark 1.1 | 9 July 2026 | 1,048,576 | 943,718 | $1.25 / $4.25 | First paid, OpenAI-compatible API; video and PDF input |
| Muse Spark 1.2 | 5 August 2026 | 1,048,576 | 943,718 | $1.25 / $4.25 | Coding point release, shipped alongside Muse Code |
| Muse Spark 1.2 Contributor | 21 August 2026 | 1,048,576 | 943,718 | $0.10 / $0.20 | Data-for-discount tier on the 1.2 base |
| Muse Spark 1.3 | 2 September 2026 | 1,048,576 | 943,718 | $1.25 / $4.25 | Long-running agentic and multi-agent work |
| Muse Spark 1.3 Contributor | 2 September 2026 | 1,048,576 | 943,718 | $0.10 / $0.20 | Contributor SKU shipped same day, not three weeks later |
Two things worth reading off that table. First, the sticker price has not moved once since the API opened in July — $1.25 and $4.25 through three generations, while the capability underneath climbed substantially. Meta is holding price and shipping capability, which is the opposite of how most vendors handle a flagship refresh.
Second, the Contributor gap closed. The 1.2 Contributor SKU arrived sixteen days after 1.2 itself; the 1.3 Contributor SKU shipped the same day as 1.3. Meta has clearly decided the cheap tier is part of the launch, not an afterthought.
Is Muse Spark still closed-weights?
Yes. As of 3 September 2026, no Muse Spark weights of any version have been published.
This matters because the position looked like it was about to change. In August, Mark Zuckerberg said Meta would open-source Muse Spark 1.2's weights. That has not happened. Meta's Hugging Face organisation currently hosts four models — Muse-Glimmer-30B, a GGUF build, an ExecuTorch build and a 3B Muse-Glimmer-30B-assistant — and no Muse Spark of any version. OpenRouter's listings for 1.1, 1.2 and 1.3 all carry a null Hugging Face id, which is the marketplace's way of saying there is nothing to download. Artificial Analysis still classifies Muse Spark 1.3 as proprietary.
So the honest state of play: Meta has stated an intention, given no date, and shipped nothing. Plan as though Muse Spark is permanently hosted-only. If open weights are a hard requirement, the model you actually want from Meta is Muse Glimmer 30B, which is genuinely Apache 2.0 and genuinely downloadable today.
What's new in Muse Spark 1.3?
Muse Spark 1.3 listed on 2 September 2026 with identical headline specs to 1.2 — 1,048,576-token context, 943,718 max output, $1.25/$4.25, the same five input modalities. The change is in what the model is tuned to do.
Meta positions 1.3 for long-running agentic, multi-agent and coding workflows, describing it as designed to keep track of information across extended tasks, work through conflicting inputs, and ask for clarification rather than guess. That is a direct answer to the specific criticism 1.2 attracted: it was fast and cheap on bounded work but drifted on long-horizon autonomous runs, where one early wrong assumption compounds for hours.
The clearest measurable signal is the composite index. On Artificial Analysis's Intelligence Index v4.3.2 (5 October 2026), Muse Spark 1.3 scores 48.09 at max reasoning effort and 45.07 at xhigh. The same basket puts 1.2 at 39.58 and 1.1 at 33.73 — about 14.4 points of composite gain across two months at a completely flat price, and 8.5 points in the four weeks between 1.2 and 1.3. We have dropped the rank position we previously quoted here: AA's measured field keeps growing, so a rank is stale within weeks even when the score behind it is right.
Two things that were unavailable at launch have since landed, and both complicate the headline. First, speed: AA now measures 1.3 at 145.9 output tokens per second with a 47.9-second time to first token at max effort, against 1.2's 270.1 tok/s and 11.7 seconds. The capability gain costs roughly half the throughput and four times the initial latency — a real trade, not a free upgrade.
Second, per-eval results. AA's v4.3.2 basket includes Terminal-Bench 4.0, and the component scores are now published — which is the first neutral agentic-coding number the Muse Spark line has had. They are much less flattering than the composite:
| Model (max effort) | Terminal-Bench 4.0 | Intelligence Index v4.3.2 |
|---|---|---|
| GPT-6.1 Sol | 56.1% | 51.83 |
| Claude Opus 5 | 49.0% | 50.78 |
| Muse Spark 1.3 | 33.3% | 48.09 |
| Muse Spark 1.3 (xhigh) | 16.7% | 45.07 |
| Muse Spark 1.2 | 7.1% | 39.58 |
Read those two columns together. On the composite, Muse Spark 1.3 sits 2.7 points behind Claude Opus 5 — close. On the one agentic-coding eval in the basket, it scores 33.3% against Opus 5's 49.0% and GPT-6.1 Sol's 56.1% — not close. The composite flatters Muse Spark because it averages in the columns where the model genuinely is frontier-grade: it ties the best models on AA's long-context reasoning eval at 0.83, and beats Opus 5 on SciCode (58.8% against 56.4%). The gap concentrates in exactly the capability the Muse line is marketed on. Treat the Terminal-Bench numbers as directional — the task count is small enough that single-digit moves are noise — but the ordering has been stable across both 1.3 effort tiers. Our Muse Spark 1.3 complete guide and the 1.3 vs 1.2 comparison track those as they land.
One documented limitation is worth knowing before you build: Meta states that audio understanding in Muse Spark 1.3 is not fully supported and that response quality on requests containing audio may be degraded. Audio is listed as an accepted input modality, but it is not production-grade yet.
How much does Muse Spark cost?
Three prices, depending on how you reach it.
- Consumers: free. Muse Spark powers Meta AI at meta.ai, in the Meta AI app, and across WhatsApp, Instagram, Facebook, Messenger and Meta's AI glasses at no charge.
- Standard API: $1.25 input / $4.25 output per million tokens, with cached input at $0.15 — an 88% cache discount, which is aggressive and rewards stable system prompts.
- Contributor tier: $0.10 input / $0.20 output, with cached input at $0.002. That is 12.5× cheaper on input and 21× cheaper on output.
Here is where standard Muse Spark 1.3 lands against the current frontier, using live list prices and Artificial Analysis Intelligence Index v4.3.2 scores captured on 5 October 2026. The last column is AA's own cost-per-Index-task figure, which is the right number for comparing models — it normalises for the fact that a more verbose model burns more tokens to reach the same answer:
| Model (max effort) | Input / 1M | Output / 1M | AA Index v4.3.2 | AA cost / Index task |
|---|---|---|---|---|
| Muse Spark 1.3 | $1.25 | $4.25 | 48.09 | $1.60 |
| Muse Spark 1.3 Contributor | $0.10 | $0.20 | same base model | ~1/13th of standard |
| GPT-6.1 Sol | $2.00 | $10.00 | 51.83 | $0.72 |
| GLM-5.3 | $1.40 | $4.40 | 44.78 | $2.01 |
| xAI Grok 4.6 | $2.00 | $6.00 | 44.31 | $1.86 |
| GPT-5.6 Sol | $2.00 | $10.00 | 46.97 | $1.99 |
| Kimi K3 | $3.00 | $15.00 | 43.59 | $2.00 |
| Claude Opus 5 | $5.00 | $25.00 | 50.78 | $5.86 |
| Claude Fable 5.1 | $10.00 | $50.00 | 53.35 | $7.63 |
The Muse Spark argument still mostly holds, with one real crack in it. Claude Opus 5 is 2.7 index points ahead and costs 3.7× more per Index task ($5.86 against $1.60). Claude Fable 5.1 is 5.3 points stronger at 4.8× the per-task cost. Grok 4.6, GPT-5.6 Sol, GLM-5.3 and Kimi K3 are all behind Muse Spark 1.3 on the Index and all cost more per task — for those four, Muse Spark wins on both axes outright.
The crack is GPT-6.1 Sol. It scores 51.83 — 3.7 points above Muse Spark 1.3 — at $0.72 per Index task, less than half Muse Spark's $1.60, despite a higher sticker price per token. It is more efficient with reasoning tokens, and on AA's normalised measure that matters more than the list rate. So the honest version of the claim: Muse Spark 1.3 is the best price-to-index model in Meta's own comparison set and against most of the near-frontier field, but it is no longer the best price-to-index model on the board. This is also the clearest illustration of why cost per task beats cost per token: GPT-6.1 Sol lists at 2.4× Muse Spark's output rate and still comes out cheaper on the same work.
The caveat that ruins naive cost models: reasoning is mandatory on Muse Spark and hidden chain-of-thought tokens bill at the $4.25 output rate. A task that thinks hard can cost several times what the visible output length suggests. Before you migrate a workload on the sticker price, run your real prompts at your intended reasoning_effort and measure cost per solved task, not per token.
What is the Contributor tier, and what does it actually cost you?
The Contributor SKU is the most interesting pricing move in the family and the one most likely to cause a problem in a company that does not read the terms.
At $0.10 input and $0.20 output it is a 12.5×/21× discount on identical model capability. The consideration is stated plainly in Meta's own listing: "Prompts and outputs may be used to improve Meta's products." You are paying with your data instead of your budget.
The governance problem is the delivery mechanism. Contributor mode is not a signed agreement or an account-level setting — it is a different model id. Any engineer who changes one string in a config file has moved that workload onto a tier where prompts and completions feed Meta's training. There is no procurement step, no approval gate, nothing that would show up in an audit until someone greps the config. If you run Muse Spark anywhere near customer data or proprietary source, the contributor model ids belong on a deny-list in your gateway. We work through the specifics in what 21× cheaper actually costs you.
How does Muse Spark relate to the rest of the Muse family?
Muse Spark is one line inside a family Meta has been building out at speed. The rest of it:
- Muse Glimmer 30B (9 August 2026) — the open-weights sibling. Roughly 29.6B parameters including a ~1.8B ViT-G/14 perception encoder, 52-layer dense transformer, 131,072-token context, text and image in, Apache 2.0, distilled from Muse Spark and tuned for autonomous agents on consumer hardware. Meta's model card reports 76.0 on SWE-Bench Verified, 51.2 on SWE-Bench Pro, 75.5 on MCP Atlas, 94.7 on AIME 2026 and 83.5 on GPQA Diamond. GGUF and ExecuTorch builds plus a 3B assistant variant ship alongside it. This is the answer whenever "I need Meta, but self-hosted" comes up.
- Muse Code (5 August 2026) — Meta's terminal coding agent, in the same category as Claude Code and Codex CLI. Approvals and an OS-enforced sandbox on by default, parallel subagents in isolated git worktrees. It is the harness Muse Spark's coding tuning was built around.
- Muse Image and Muse Video — Meta's generative media models, announced 7 July 2026. Muse Image reached OpenRouter on 26 August 2026 at $0.01 per image.
- Muse Voice Transcribe (2 September 2026) — a real-time speech model shipped the same day as Muse Spark 1.3, at $0.18 per hour of audio ($3 per 1,000 minutes) with speaker diarization for 20+ speakers, 80ms streaming chunks, 70+ languages, a reported 3.1% word error rate and 17.5% average diarization error rate.
The pattern across all of it: Meta is not trying to own the top of any leaderboard. It is trying to be the cheapest credible option in every category at once — reasoning, open weights, coding agent, image, speech. For how that reshapes the wider field, see our open-source LLMs landscape pillar and the Llama 4 guide this family superseded.
How does Muse Spark compare with Claude, GPT and Grok?
The composite indexes above tell you the ranking; they do not tell you where each model actually wins. The practical split, based on what has been independently measured so far:
- vs Claude Fable 5.1 and Claude Opus 5 — Anthropic still leads on hard single-shot coding correctness and on long-horizon reliability. Muse Spark's case is that it is 2.7 to 5.3 index points behind on AA's v4.3.2 composite at a quarter to a fifth of the cost per Index task — though the gap on Terminal-Bench 4.0 specifically is far wider (33.3% against Opus 5's 49.0%), so the composite understates the distance on agentic coding. If a wrong answer is expensive, pay for Claude; if you are running the same bounded task ten thousand times, the arithmetic flips hard.
- vs GPT-5.6 Sol — Sol is behind Muse Spark 1.3 on AA v4.3.2 (46.97 against 48.09) and costs $10 per million output against $4.25, so Muse Spark wins on both axes. That stops being true one generation up: GPT-6.1 Sol scores 51.83 at $0.72 per Index task against Muse Spark's $1.60, and leads Terminal-Bench 4.0 at 56.1%. If you are comparing against OpenAI, compare against 6.1, not 5.6.
- vs Grok 4.6 — behind on AA v4.3.2 (44.31 at high effort against 48.09) and more expensive per Index task ($1.86 against $1.60), but the one to check pricing on carefully: Grok 4.6 lists $2/$6 with a 500K context, but bills the entire request at a higher rate above 200K tokens. Muse Spark's 1M context has no such cliff.
- vs GLM-5.3 — the closest price competitor at $1.40/$4.40, but no longer close on capability: 44.78 at max effort against Muse Spark 1.3's 48.09, at $2.01 per Index task against $1.60. The GLM line's real pressure comes from GLM 5.3 Flash, which scores 41.81 for $0.25 per task — well behind on capability, far ahead on cost.
The one gap worth naming honestly: Muse Spark performs better on composite and preference boards than on agentic-coding evals specifically. It was long absent from the verified agentic-coding leaderboards — Terminal-Bench, SWE-bench, SWE-rebench, Aider — and the first neutral number to arrive confirms the suspicion rather than dispelling it. Artificial Analysis's v4.3.2 basket includes Terminal-Bench 4.0, where Muse Spark 1.3 scores 33.3% against Claude Opus 5's 49.0% and GPT-6.1 Sol's 56.1% — a far bigger gap than the 2.7-point composite difference implies. That divergence is the central question about the whole line, and we take it apart with the 1.2 numbers in Muse Spark 1.2 benchmarks vs Claude Opus 5, and against the current version in Muse Spark 1.3 vs Claude Opus 5.
Which Muse Spark version should you use today?
Simple, because Meta has kept the pricing flat and the specs identical:
- Use 1.3 for everything new, with one exception. Same context, same max output, same price as 1.2, and 8.5 index points higher on AA v4.3.2 (48.09 against 39.58) with explicit tuning for long-running agentic work. There is no cost argument for staying on 1.2 — but there is a latency argument: 1.2 runs at 270.1 tok/s with an 11.7-second time to first token against 1.3's 145.9 tok/s and 47.9 seconds. For interactive and streaming workloads, 1.2 is still the better experience. Note too that 1.3 costs $1.60 per AA Index task against 1.2's $0.97 at identical token rates, purely from extra reasoning tokens.
- Use the Contributor SKU only for genuinely non-sensitive work — open-source repos, throwaway experiments, public data, learning. Never for customer data or proprietary source, and put the contributor model ids behind a gateway deny-list so the choice cannot be made accidentally.
- Use Muse Glimmer 30B if you need weights you control. It is the only Meta model in this family you can actually download.
- Avoid Muse Spark 1.3 for audio-heavy pipelines until Meta lifts the degraded-quality caveat. If speech is the job, Muse Voice Transcribe is the purpose-built model.
What are the risks of building on Muse Spark?
- No exit. Closed weights, no self-hosting, no fine-tuning, and an open-weights promise Meta has made but not kept. Anything you build on Muse Spark is portable only as far as your prompts are.
- Benchmark provenance. Meta drew benchmark-gaming criticism during the Llama 4 / LMArena episode, and its 1.2 numbers were published as chart images without a released harness config or raw results. Weight independent harnesses over vendor charts, and weight verified leaderboards over both.
- Data terms drift by model id. The Contributor tier means the privacy properties of your workload depend on a string in a config file. That is an unusual failure mode and worth an explicit control.
- Release cadence. Four versions in five months is great for capability and hard for reproducibility. Pin an exact dated model id — the listings expose them, for example
meta/muse-spark-1.3-20260902— rather than a floating alias, or your evals will drift underneath you.
FAQ
What is the latest version of Muse Spark?
Muse Spark 1.3, listed 2 September 2026. It carries a 1,048,576-token context window, 943,718 max output tokens, and costs $1.25 per million input tokens and $4.25 per million output tokens — identical pricing and specs to 1.2, with tuning aimed at long-running agentic and multi-agent workflows.
Is Muse Spark open source?
No. No Muse Spark weights have been released for any version. Mark Zuckerberg said in August 2026 that Meta would open-source Muse Spark 1.2's weights, but as of 3 September 2026 nothing has shipped and no date has been given. Meta's open-weights model is Muse Glimmer 30B, released under Apache 2.0.
How much does the Muse Spark API cost?
$1.25 per million input tokens, $4.25 per million output tokens, and $0.15 per million cached input tokens — an 88% cache discount. A Contributor tier costs $0.10 and $0.20 in exchange for Meta training on your prompts and outputs. Consumer access through Meta AI is free.
What is the difference between Muse Spark and Muse Glimmer?
Muse Spark is Meta's closed, hosted frontier reasoning line with a 1M-token context and five input modalities. Muse Glimmer 30B is a ~29.6B-parameter open-weights model distilled from Muse Spark, released under Apache 2.0 with a 131K context and text-plus-image input, built to run agents on consumer hardware.
Can you turn off reasoning on Muse Spark?
No. Reasoning is mandatory on every Muse Spark SKU. You can set reasoning_effort to minimal, low, medium, high or xhigh — medium is the default — but you cannot disable it. Because hidden reasoning tokens bill at the output rate, effort level is effectively a cost control.
What is Muse Spark's context window?
1,048,576 tokens (1M) on versions 1.1, 1.2 and 1.3, with maximum output of 943,718 tokens. The original consumer-only 1.0 was roughly 262K. Unlike some rivals, Muse Spark does not change its billing rate for very large prompts.
Is Muse Spark good at coding?
It is competitive rather than leading, and the gap is wider on coding than on general reasoning. Meta's own 1.2 launch charts placed it second to Claude Opus 5 on all three coding benchmarks it published. Artificial Analysis's Intelligence Index v4.3.2 now runs Terminal-Bench 4.0 as a component, and Muse Spark 1.3 scores 33.3% there against Claude Opus 5's 49.0% and GPT-6.1 Sol's 56.1% — a much larger gap than the 2.7-point composite difference suggests. Its strength is coding throughput per dollar, not top-of-leaderboard correctness.
Which Muse Spark model id should I use in production?
Pin a dated id such as meta/muse-spark-1.3-20260902 rather than a floating alias, so a mid-quarter model swap cannot silently change your eval results. Keep the -contributor ids off any production allow-list unless you have explicitly accepted that prompts and outputs may train Meta's models.
So should you build on Muse Spark?
Muse Spark's argument has changed shape. At 1.1 it was "near-frontier for a quarter of the price." At 1.3 it is 2.7 points behind Claude Opus 5 on Artificial Analysis's Intelligence Index v4.3.2 at roughly a quarter of the cost per Index task, with the sticker unmoved through three generations — a genuinely strong position. Two things keep it from being the obvious default. GPT-6.1 Sol now beats it on both capability and AA's cost-per-task at once. And on Terminal-Bench 4.0, the one agentic-coding eval in the current basket, the gap to the frontier is 15 points rather than 3.
The decision rule that holds: pick Muse Spark when volume is high, tasks are verifiable, and a wrong answer is cheap to catch. Pick Claude or GPT when a single failure is expensive. Pick Muse Glimmer when you need the weights. And whichever you pick, run your own eval set on your own prompts before migrating — the gap between second and fourth place on a leaderboard is smaller than the gap between two prompt strategies on your codebase.
Building on models that ship four versions in five months, and need engineers who already evaluate and migrate this way? Codersera helps you hire vetted remote developers fluent in agentic AI tooling — start with a risk-free trial.