Meta launches Muse Spark 1.3, aimed at more autonomous agents
Meta updates its agentic and coding model with better long-instruction handling, a lower cost per task than rivals, and the top spot on a banking benchmark, according to Artificial Analysis.

Meta has announced Muse Spark 1.3, a new version of its AI model built for agentic and coding tasks. The model is already available on Muse Code and through the Meta Model API, the interface that exposes the company's models to outside developers. Meta says this version was designed to handle long, multi-step tasks more reliably: given an open-ended goal, the model uses tools to gather information, identifies gaps in its own planning, and keeps track of what it learns along the way.
How the model handles what it doesn't know
When an instruction is unclear, Muse Spark 1.3 can pause its work to ask for clarification rather than proceed on a guess. Faced with an obstacle it cannot get past on its own, it can ask for help; and before taking an action with significant consequences, it asks for confirmation. Users can also choose between frequent progress updates or letting the system work in the background until it finishes.
The gains Meta is highlighting
The company lists several improvements over the previous version: greater accuracy in following long instructions, better preservation of requirements and constraints throughout a task, improved handling of multiple requests within a single conversation, better awareness of what the model does and doesn't know, and a reduced tendency to fabricate results. On coding tasks, Meta says the model is now less verbose and more efficient.
- About 20% fewer tool calls than Muse Spark 1.2, in internal comparative tests
- About 25% fewer tokens used in the same comparison
- The maximum reasoning mode ('max') will only ship after additional safety testing, with no date given by Meta
The fourth model in the series in five months
Muse Spark 1.3 is the series' fourth release in five months: the line launched in April, followed by 1.1 in July and 1.2 in August, according to German outlet The Decoder, which cross-referenced the figures with Artificial Analysis benchmarks. The new version ships in two tiers: 'max', in a limited partner preview, and 'xhigh', already generally available. On Artificial Analysis's Intelligence Index, the max tier scores 62 and xhigh scores 61, up from 57 in August and 53 in July.
What changed on agent benchmarks
On τ³-Bench Banking, a test that evaluates agents handling tools in a simulated banking scenario, the max tier reaches 52%, currently the top score on Artificial Analysis's leaderboard; xhigh reaches 47%, tied with Anthropic's Claude Fable 5.1 (max) and with GLM-5.3-Flash. The previous version, 1.2, scored 35% on the same test. On Terminal-Bench 2.1, focused on terminal coding, xhigh climbed from 80% to 85% and max reached 86% — still behind Claude Fable 5.1, which leads the table with 91.4% at its max tier.
On GDPval-AA v2, a set of 220 real-world professional tasks scored on an Elo scale calibrated against human performance, Meta moved from 1615 to 1709 and then to 1754 points across versions. Claude Fable 5.1 (max) remains ahead, at 1853 points; to close part of that gap, Muse Spark's max variant uses 62% more reasoning tokens than xhigh. On GPQA Diamond, a battery of expert-level science questions, Muse Spark went from 90% to 94% — in the top tier, but behind Gemini 3.8 Flash (high, 95.3%) and Grok 4.6 (high, 94.9%). On CritPt, a research-physics test, the score rose from 18% to 26%, with the lead still held by GPT-5.6 Sol (max, 32.3%) and Claude Fable 5.1 (xhigh, 31.1%).
Not everything improved: AA-LCR fell from 83% to 79% compared with version 1.2, and factual accuracy on AA-Omniscience dropped by up to three points, as the model now refuses to answer more often when it is uncertain.
Price unchanged, security tightened
On pricing, Meta kept $1.25 per million input tokens and $4.25 per million output tokens — unchanged from the previous version. An Intelligence Index task now costs about $0.55 on average, the lowest among models scoring 59 or above on the index; rivals charge between $0.94 and $1.23 for the same task. Muse Spark 1.3 is still pricier than 1.2, which cost $0.40 per task — Meta has not yet announced pricing for the max variant. The company also said larger models are coming, along with a future open-weights release.
On safety, Meta says Muse Spark 1.3 is more resistant to adversarial inputs and prompt-injection attempts, and does a better job assessing whether an action is irreversible before carrying it out. That's why the maximum reasoning mode will only open up after additional safety testing — Meta gave no timeline.
For developers and startups weighing Meta's more open ecosystem against closed models from Anthropic or Google for agentic coding, Muse Spark 1.3 shifts the calculation in two ways. First, a per-task cost among the lowest of any frontier-class model lowers the bill for long-running agents — a real saving for teams running agentic workloads at scale. Second, Meta's confirmation that an open-weights version is coming strengthens an option that startups wary of vendor lock-in and API pricing volatility have been watching closely: the ability to run the model on their own infrastructure instead of depending on a closed API. The gap hasn't closed — Claude Fable 5.1 still leads on most of the benchmarks cited here — but it has narrowed enough that the choice is no longer automatic.
Sources
- Meta lança Muse Spark 1.3 com foco em agentes mais autônomosOlhar Digital · September 3, 2026
- Meta schließt mit Muse Spark 1.3 zur Spitze auf, lockt mit niedrigen PreisenThe Decoder · September 2, 2026



