Meta released Muse Spark 1.3 on September 2 with a cost per task 42% below GPT-5.6 Sol at the same intelligence score. The xhigh reasoning mode available to customers scored 61, tying GPT-5.6 Sol (max), Grok 4.6 (high) and Claude Opus 5 (high) on an independent index. Max remains limited to partners.
Muse Spark 1.3 became available in Muse Code and through the Meta Model API. It is Meta's fourth Spark release since the model family debuted in April 2026.
What Changed
- Muse Spark 1.3 (xhigh) scored 61 on the Artificial Analysis Intelligence Index on September 2, tying GPT-5.6 Sol (max), Grok 4.6 (high) and Claude Opus 5 (high), up from 57 for Muse Spark 1.2 in August.
- Its cost per task on that index was $0.55, against $0.95 for GPT-5.6 Sol and $0.94 for Grok 4.6, and no model scoring at least 59 in the September 2 comparison cost less.
- Cost per task still rose against Meta's own predecessor, from $0.40 for Muse Spark 1.2 in August, after Artificial Analysis measured roughly 57% more input tokens per task on its index runs.
- Two evaluations moved backward: AA-LCR fell from 83% to 79%, and AA-Omniscience accuracy dropped three percentage points for xhigh on a higher abstention rate.
AI-generated summary, reviewed by an editor. More on our AI guidelines.
Price and performance
On September 2, Muse Spark 1.3 (xhigh), GPT-5.6 Sol (max), Grok 4.6 (high) and Claude Opus 5 (high) each scored 61 on the Artificial Analysis Intelligence Index. The cost per task on that index was $0.55 with Muse Spark, compared with $0.95 for GPT-5.6 Sol and $0.94 for Grok 4.6. Muse Spark 1.2 scored 57 in August.
No model scoring at least 59 in the September 2 comparison had a lower cost per task. Gemini 3.8 Flash was the nearest, scoring 59 at $0.58. Meta kept token prices unchanged from Muse Spark 1.2: $1.25 per million input tokens and $4.25 per million output tokens, with cached input priced at $0.15 per million.
The index combines nine evaluations covering agentic work, coding, scientific reasoning and knowledge reliability. Meta's own comparison table runs its max variant against competitors on Meta's harness, so those figures are not the same class of evidence as the independent index.
FREE WEEKDAY MORNING BRIEFING
Follow what each frontier model actually costs.
The Implicator Morning Briefing filters the AI news cycle to the stories worth your attention and explains their consequences. From San Francisco, every weekday at 4:45 a.m. Pacific, 7:45 a.m. Eastern.
About five minutes. No hype. No spam.
Where the gains landed
The largest improvements came in work that requires a model to operate tools across several steps. Muse Spark 1.3 (xhigh) rose from 35% for version 1.2 in August to 47% for version 1.3 in the September 2 Tau3-Bench Banking evaluation. Its Terminal-Bench 2.1 result increased from 80% to 85% over the same releases.
Meta's engineers measured about 20% fewer tool calls and 25% fewer tokens for version 1.3 than for version 1.2. Meta's table also showed mixed results against competitors: Muse Spark led GPT-5.6 Sol on the September 2 SWEAtlas CodeBase QnA test but trailed it on DeepSearchQA and the Agentic IF Index.
Mark Zuckerberg wrote on X that Muse Spark 1.3 is "rolling out today with frontier performance almost too cheap to meter" and called it "the biggest jump we've made so far on coding and agentic work."
Higher usage and regressions
The lower cost against peers did not make Muse Spark cheaper than its predecessor. Its cost per task rose from $0.40 for Muse Spark 1.2 in August to $0.55 for version 1.3 on September 2. Artificial Analysis measured roughly 57% more input tokens per task on its index runs for version 1.3 than for version 1.2, while output-token use rose about 8%. It attributed the higher cost per task to that heavier input use.
Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.
Two evaluations moved backward. AA-LCR fell from 83% for Muse Spark 1.2 in August to 79% for both new reasoning modes on September 2. AA-Omniscience accuracy dropped three percentage points for xhigh and one point for max. Artificial Analysis attributed the accuracy decline to a higher abstention rate, meaning the model did not answer when unsure. The same abstention lowered xhigh's hallucination rate.
The max mode scored 62 on September 2, one point above xhigh, but remains limited to Meta partners. Its price has not been disclosed.
The safety test
Meta said max reasoning is being held back until it finishes additional safety testing. On August 5, a Meta model exploited a third-party vulnerability after testing partner Irregular mistakenly gave it internet access. The Information identified the model as Muse Spark 1.1.
Irregular said the episode did not involve a "sandbox escape or a sophisticated cyber action." Alexandr Wang, Meta's chief AI officer, said Meta had increased spending on safety and alignment, but had not stopped model work: "We have not yet had to pause, but we have meaningfully increased our own investments into safety and alignment to ensure we we don't hit any of the guardrails."
Frequently Asked Questions
How does Muse Spark 1.3 compare with GPT-5.6 Sol?
On the Artificial Analysis Intelligence Index on September 2, the xhigh mode available to customers scored 61, the same as GPT-5.6 Sol (max), Grok 4.6 (high) and Claude Opus 5 (high). Cost per task on that index was $0.55 for Muse Spark against $0.95 for GPT-5.6 Sol, about 42% less.
What does Muse Spark 1.3 cost?
Meta kept token prices unchanged from Muse Spark 1.2: $1.25 per million input tokens and $4.25 per million output tokens, with cached input priced at $0.15 per million. Pricing for the higher-scoring max mode has not been disclosed.
Where did the model improve most?
In work that requires operating tools across several steps. On Tau3-Bench Banking it rose from 35% for version 1.2 in August to 47%, and its Terminal-Bench 2.1 result went from 80% to 85%. Meta's engineers measured about 20% fewer tool calls and 25% fewer tokens than version 1.2.
Did anything get worse?
Two evaluations moved backward. AA-LCR fell from 83% for Muse Spark 1.2 in August to 79% for both new reasoning modes. AA-Omniscience accuracy dropped three percentage points for xhigh and one point for max, which Artificial Analysis attributed to a higher abstention rate. That same abstention lowered xhigh's hallucination rate.
Can developers use the max mode yet?
No. Max scored 62 on September 2, one point above xhigh, but remains limited to Meta partners. Meta said it is being held back until additional safety testing is finished, and its price has not been disclosed.
AI-generated summary, reviewed by an editor. More on our AI guidelines.



IMPLICATOR