
Muse Spark 1.3 Hits the Frontier at $0.55 per Task
Meta's Muse Spark 1.3 scores 61 on the Artificial Analysis Intelligence Index at $0.55 per task, the cheapest model measured above a score of 59.
Meta Closes the Gap With a Post-Training Update
Meta released Muse Spark 1.3 on September 2, 2026, and the headline number is not the intelligence score. It is the price attached to it. Independent evaluator Artificial Analysis puts the model's xhigh reasoning setting at 61 on its Intelligence Index, level with GPT-5.6 Sol at max reasoning and Grok 4.6 at high reasoning, and it costs $0.55 to run the index's full task set. GPT-5.6 Sol at max costs $0.95 for the same work. That makes Muse Spark 1.3 the cheapest model Artificial Analysis has measured at a score of 59 or above.
- Intelligence Index: 61 at xhigh reasoning, up 4 points from Muse Spark 1.2 in August and 8 points from 1.1 in July
- Cost: $0.55 per Intelligence Index run, against $0.95 for GPT-5.6 Sol at max reasoning
- Pricing unchanged: $1.25 per million input tokens and $4.25 per million output tokens, with a 1M-token context window
- Efficiency: Meta reports roughly 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2 on its internal coding work
What Actually Improved in Muse Spark 1.3
The gains are concentrated rather than spread thin, which is what you would expect from a post-training update rather than a new pretraining run. Artificial Analysis attributes most of the movement to two areas. On Tau3-Bench Banking, an agentic benchmark that measures whether a model can complete multi-step customer workflows without losing the thread, the score moved from 35% to 47%. On CritPt, a scientific reasoning set, it went from 18% to 26%.
The rest of the index moved less, and two components moved backwards. Artificial Analysis records a 4-point drop on AA-LCR, its long-context reasoning test, and a 1 to 3 point decline in AA-Omniscience accuracy. This is the honest shape of a targeted update: you tune for agentic reliability and coding, and something at the edges gives a little ground. It is worth knowing before you swap a production long-context pipeline over.
Meta's own account of the release emphasises behaviour rather than benchmarks. The company says the model sustains longer-horizon work across multiple workflows in a single thread, asks clarifying questions on ambiguous requests instead of guessing, and shows better awareness of what it does not know. Meta also describes stronger defences against adversarial inputs and improved judgement about irreversible actions in agentic settings — the kind of thing that matters far more than a leaderboard position once an agent has write access to anything.
Why Does Cost per Task Matter More Than Cost per Token?
Token pricing is the number everyone quotes, and it is close to useless for comparing reasoning models. A model that thinks for 4,000 tokens before answering at $1 per million can easily cost more per completed job than one that thinks for 900 tokens at $3 per million. Cost per task folds the reasoning effort back in.
That is why the $0.55 figure is the interesting part of this release. Meta did not cut its list price — input and output rates are identical to Muse Spark 1.2. It got more index points out of the same token budget, and the efficiency claim it makes about its own engineering work, roughly 20% fewer tool calls and 25% fewer tokens, points at the same mechanism. For anyone running agents at volume, that arithmetic decides the monthly bill more than the rate card does.
Where Does Muse Spark 1.3 Sit Among Frontier Models?
Artificial Analysis is measuring two variants. The xhigh setting, generally available now, sits at 61. A max variant, currently in limited preview for Meta partners, scores 62 — which on that index places it behind only Anthropic's Fable and Opus lines. Meta itself frames the release as competitive with GPT-5.6 Sol and Opus 5 across agentic, coding, instruction-following and long-context evaluations.
Treat the vendor comparisons as vendor comparisons. The independent number is the index score, and 61 puts Meta genuinely in the frontier band for the first time on a generally available model rather than a preview. It is also Meta's fourth iteration of this model line in five months, following 1.1 in July and 1.2 in August, which is a release cadence worth noticing on its own.
Muse Spark 1.3 handles text, image and video input, and is available through Meta's first-party API and through Muse Code, the terminal agent we covered when Muse Code shipped alongside Muse Spark 1.2 in August.
What This Means If You Build on These Models
The practical read is that the frontier band now has four or five occupants rather than two, and they are converging on price. A month ago the choice between a top-tier agentic model and a cheap one carried a real capability cost. On this week's numbers it carries much less, and the gap that remains is narrowest exactly where most production work happens — tool use, multi-step tasks and code.
That is good news for anyone building on top of these systems, and it is the same trend visible in Anthropic's Fable 5.1 release last week and across our AI coverage this month: capability improvements are arriving as efficiency improvements, and the models are getting cheaper to actually finish a job with.
Sources: Meta AI Research — Introducing Muse Spark 1.3 — September 2, 2026; Artificial Analysis — Muse Spark 1.3: Meta reaches the frontier — September 2026; Axios — September 2, 2026.
More AI Stories

Commerce Agent Blueprints Let Retailers Ship in Days
Anthropic's commerce agent blueprint ships reference code for shopping and merchant agents, with early users reporting carts up to 35% larger.

K2 Horizon Ships Six Open Models With Training Data
IFM's K2 Horizon releases six Apache 2.0 models from 0.9B to 375B parameters, publishing training data, code and logs alongside the weights.

Wayve AI Driver Brings Robotaxi Rides to London Uber
Wayve and Uber launched the UK's first autonomous rides in London on September 3, with a safety driver aboard and no surcharge on UberX fares.
