> ## Content Index
> Fetch the complete content index at: https://www.implicator.ai/llms.txt
> Use this file to discover other available public pages before exploring further.

# Meta's Muse Spark 1.3 Matches GPT-5.6 Sol at 42% Lower Cost per Task
- URL: https://www.implicator.ai/meta-muse-spark-1-3-cost-per-task/
- Published: 2026-09-03T00:27:48.000Z
- Updated: 2026-09-03T00:27:48.000Z
- Description: Meta's Muse Spark 1.3 tied GPT-5.6 Sol at 61 on the Artificial Analysis index while costing $0.55 per task against $0.95. But its cost per task rose from $0.40 for version 1.2, two evaluations went backward, and the max mode stays limited to partners with no published price.
- Author: Marcus Schuler
- Tags: AI News

Meta [released Muse Spark 1.3](https://research.meta.ai/blog/introducing-muse-spark-1-3?ref=implicator.ai) on September 2 with a cost per task 42% below GPT-5.6 Sol at the same intelligence score. The xhigh reasoning mode available to customers scored 61, tying GPT-5.6 Sol (max), Grok 4.6 (high) and Claude Opus 5 (high) on an [independent index](https://artificialanalysis.ai/models/muse-spark-1-3?ref=implicator.ai). Max remains limited to partners.

Muse Spark 1.3 became available in Muse Code and through the Meta Model API. It is Meta's fourth Spark release since the model family debuted in April 2026.

What Changed

- Muse Spark 1.3 (xhigh) scored 61 on the Artificial Analysis Intelligence Index on September 2, tying GPT-5.6 Sol (max), Grok 4.6 (high) and Claude Opus 5 (high), up from 57 for Muse Spark 1.2 in August.
- Its cost per task on that index was $0.55, against $0.95 for GPT-5.6 Sol and $0.94 for Grok 4.6, and no model scoring at least 59 in the September 2 comparison cost less.
- Cost per task still rose against Meta's own predecessor, from $0.40 for Muse Spark 1.2 in August, after Artificial Analysis measured roughly 57% more input tokens per task on its index runs.
- Two evaluations moved backward: AA-LCR fell from 83% to 79%, and AA-Omniscience accuracy dropped three percentage points for xhigh on a higher abstention rate.

AI-generated summary, reviewed by an editor. [More on our AI guidelines](https://www.implicator.ai/about/).

## Price and performance

On September 2, Muse Spark 1.3 (xhigh), GPT-5.6 Sol (max), Grok 4.6 (high) and Claude Opus 5 (high) each scored 61 on the Artificial Analysis Intelligence Index. The cost per task on that index was $0.55 with Muse Spark, compared with $0.95 for GPT-5.6 Sol and $0.94 for Grok 4.6\. Muse Spark 1.2 scored 57 in August.

No model scoring at least 59 in the September 2 comparison had a lower cost per task. Gemini 3.8 Flash was the nearest, scoring 59 at $0.58\. Meta kept token prices unchanged from Muse Spark 1.2: $1.25 per million input tokens and $4.25 per million output tokens, with cached input priced at $0.15 per million.

The index combines nine evaluations covering agentic work, coding, scientific reasoning and knowledge reliability. Meta's own comparison table runs its max variant against competitors on Meta's harness, so those figures are not the same class of evidence as the independent index.

FREE WEEKDAY MORNING BRIEFING

Follow what each frontier model actually costs.

The Implicator Morning Briefing filters the AI news cycle to the stories worth your attention and explains their consequences. From San Francisco, every weekday at 4:45 a.m. Pacific, 7:45 a.m. Eastern.

Email address 

Send me tomorrow’s briefing 

Check your inbox for the confirmation link.

About five minutes. No hype. No spam.

## Where the gains landed

The largest improvements came in work that requires a model to operate tools across several steps. Muse Spark 1.3 (xhigh) rose from 35% for version 1.2 in August to 47% for version 1.3 in the September 2 Tau3-Bench Banking evaluation. Its Terminal-Bench 2.1 result increased from 80% to 85% over the same releases.

Meta's engineers measured about 20% fewer tool calls and 25% fewer tokens for version 1.3 than for version 1.2\. Meta's table also showed mixed results against competitors: Muse Spark led GPT-5.6 Sol on the September 2 SWEAtlas CodeBase QnA test but trailed it on DeepSearchQA and the Agentic IF Index.

[Mark Zuckerberg wrote on X](https://x.com/finkd/status/2095232032896946311?ref=implicator.ai) that Muse Spark 1.3 is "rolling out today with frontier performance almost too cheap to meter" and called it "the biggest jump we've made so far on coding and agentic work."

## Higher usage and regressions

The lower cost against peers did not make Muse Spark cheaper than its predecessor. Its cost per task rose from $0.40 for Muse Spark 1.2 in August to $0.55 for version 1.3 on September 2\. Artificial Analysis measured roughly 57% more input tokens per task on its index runs for version 1.3 than for version 1.2, while output-token use rose about 8%. It attributed the higher cost per task to that heavier input use.

Know someone who'd find this useful? [✉️ Email it to a friend in one click](mailto:?subject=A%20newsletter%20I%20think%20you%27d%20like&body=This%20is%20one%20of%20maybe%20three%20newsletters%20I%20actually%20read.%20The%20rest%20just%20pile%20up%2C%20unread%2C%20judging%20me.%0A%0AAnd%20yes%2C%20this%20email%20mostly%20wrote%20itself%2C%20which%20is%20a%20little%20on%20the%20nose%20for%20an%20AI%20newsletter.%20Doesn%27t%20make%20it%20wrong.%20implicator.ai%20is%20good.%0A%0ASubscribe%20free%3A%20https%3A%2F%2Fwww.implicator.ai%2Fsubscribe%2F%3Futm%5Fsource%3Dnewsletter%26utm%5Fmedium%3Dforward%26utm%5Fcampaign%3Demail%5Fforward), or they can [subscribe free here](https://www.implicator.ai/subscribe/?utm%5Fsource=newsletter&utm%5Fmedium=forward&utm%5Fcampaign=forward%5Fto%5Fcolleague).

Two evaluations moved backward. AA-LCR fell from 83% for Muse Spark 1.2 in August to 79% for both new reasoning modes on September 2\. AA-Omniscience accuracy dropped three percentage points for xhigh and one point for max. Artificial Analysis attributed the accuracy decline to a higher abstention rate, meaning the model did not answer when unsure. The same abstention lowered xhigh's hallucination rate.

The max mode scored 62 on September 2, one point above xhigh, but remains limited to Meta partners. Its price has not been disclosed.

## The safety test

Meta said max reasoning is being held back until it finishes additional safety testing. On [August 5, a Meta model exploited a third-party vulnerability](https://www.theguardian.com/technology/2026/aug/05/meta-ai-model-hack-training?ref=implicator.ai) after testing partner Irregular mistakenly gave it internet access. The Information identified the model as Muse Spark 1.1.

Irregular said the episode did not involve a "sandbox escape or a sophisticated cyber action." [Alexandr Wang, Meta's chief AI officer, said](https://www.axios.com/2026/09/02/meta-debuts-muse-spark-13-as-personal-agent-work-continues?ref=implicator.ai) Meta had increased spending on safety and alignment, but had not stopped model work: "We have not yet had to pause, but we have meaningfully increased our own investments into safety and alignment to ensure we we don't hit any of the guardrails."

Frequently Asked Questions

How does Muse Spark 1.3 compare with GPT-5.6 Sol?

On the Artificial Analysis Intelligence Index on September 2, the xhigh mode available to customers scored 61, the same as GPT-5.6 Sol (max), Grok 4.6 (high) and Claude Opus 5 (high). Cost per task on that index was $0.55 for Muse Spark against $0.95 for GPT-5.6 Sol, about 42% less.

What does Muse Spark 1.3 cost?

Meta kept token prices unchanged from Muse Spark 1.2: $1.25 per million input tokens and $4.25 per million output tokens, with cached input priced at $0.15 per million. Pricing for the higher-scoring max mode has not been disclosed.

Where did the model improve most?

In work that requires operating tools across several steps. On Tau3-Bench Banking it rose from 35% for version 1.2 in August to 47%, and its Terminal-Bench 2.1 result went from 80% to 85%. Meta's engineers measured about 20% fewer tool calls and 25% fewer tokens than version 1.2.

Did anything get worse?

Two evaluations moved backward. AA-LCR fell from 83% for Muse Spark 1.2 in August to 79% for both new reasoning modes. AA-Omniscience accuracy dropped three percentage points for xhigh and one point for max, which Artificial Analysis attributed to a higher abstention rate. That same abstention lowered xhigh's hallucination rate.

Can developers use the max mode yet?

No. Max scored 62 on September 2, one point above xhigh, but remains limited to Meta partners. Meta said it is being held back until additional safety testing is finished, and its price has not been disclosed.

AI-generated summary, reviewed by an editor. [More on our AI guidelines](https://www.implicator.ai/about/).

[Sonnet 5 Closes Most of the Gap to Opus 4.8 on Agent WorkAnthropic released Claude Sonnet 5 on June 30 with a simple pitch: most of Opus 4.8's capability at well under half the cost. On the company's own agent benchmarks, the model trails its flagship by a The Implicator![](https://www.implicator.ai/content/images/2026/07/2026-06-30-17.33.32-sonnet5_scale@2x.webp)](https://www.implicator.ai/sonnet-5-closes-most-of-the-gap-to-opus-4-8-on-agent-work/)

[Alibaba Ships Qwen3.6-27B, an Open-Weight Coding Model That Beats Its 397B MoEAlibaba on Wednesday released Qwen3.6-27B, a dense 27-billion-parameter open-weight model under Apache 2.0 that tops its own 397B-parameter predecessor on every major agentic coding benchmark. The modThe Implicator![](https://www.implicator.ai/content/images/2026/04/2026-04-22-17.37.35-qwen-27b-dense@2x.webp)](https://www.implicator.ai/alibaba-ships-qwen3-6-27b-an-open-weight-coding-model-that-beats-its-397b-moe/)

[Kimi K2.6 did not release a coding model. It opened the control room.On Monday, Moonshot AI put a familiar label on a less familiar move. Kimi K2.6 arrived as an open-source coding model, with a benchmark table, a Hugging Face page, a coding CLI, and the usual claims aThe Implicator![](https://www.implicator.ai/content/images/2026/04/2026-04-20-10.01.59-kimi_k2_6@2x.webp)](https://www.implicator.ai/kimi-k2-6-did-not-release-a-coding-model-it-opened-the-control-room/)