Meta closes in on the top with Muse Spark 1.3, and undercuts rivals on price

Meta has released Muse Spark 1.3, its fourth model in five months. Independent testing shows solid gains on agentic tasks, but a gap to the top stays. What sells the model is the price. Meta has released Muse Spark 1.3 through Muse Code and the Meta Model API. The series launched in April, with ...
Meta has released Muse Spark 1.3, its fourth model in five months. Independent testing shows solid gains on agentic tasks, but a gap to the top stays. What sells the model is the price.
Meta has released Muse Spark 1.3 through Muse Code and the Meta Model API. The series launched in April, with version 1.1 following in July and 1.2 in August. The xhigh tier is available now, while the more compute-heavy max tier arrives only after further safety testing and currently runs as a limited partner preview, based on to Artificial Analysis.
At an unchanged $1.25 and $4.25 per million input and output tokens, one index task costs $0.55. No model scoring 59 points or higher is cheaper, and rivals at the same index level run between $0.94 and $1.23. Muse Spark 1.3 does cost more than version 1.2, which ran $0.40.
In a related development, on the Intelligence Index, max scores 62 points and xhigh 61, up from 57 in August and 53 in July. The jump comes down to how the index weights its tests. GDPval-AA v2 counts for 20 percent, Terminal-Bench 2.1 for 16 percent, and τ³-Bench Banking for 14 percent, and Meta's biggest gains land in exactly those three tests.
On τ³-Banking, where agents operate tools in a simulated banking scenario, max hits 52 percent. That's number one right now, based on to Artificial Analysis, and it's the only outright lead the model holds. The available xhigh tier reaches 47 percent, tying Claude Fable 5.1 (max) and GLM-5.3-Flash rather than leading. The predecessor 1.2 sat at 35 percent.
Terminal-Bench 2.1, which tests coding in the terminal, climbs from 80 to 85 percent on xhigh and 86 on max, but Claude Fable 5.1 still holds the top spot at 91.4 percent in its max tier, 91.0 at xhigh, and 89.9 at high. On the index's highest-weighted test, GDPval-AA v2, Meta improves from 1,615 to 1,709 and 1,754 on a scale calibrated to human expert performance at 1,000 across 220 real-world professional tasks. Claude Fable 5.1 (max) sits at 1,853. Meta buys the max variant's edge with compute, burning 62 percent more reasoning tokens than xhigh.
Furthermore, on GPQA Diamond, which poses expert-level science questions, Muse Spark rises from 90 to 94 percent. That's the top group, but still below Gemini 3.8 Flash (high) at 95.3 and Grok 4.6 (high) at 94.9. CritPt, covering research physics, jumps from 18 to 26 percent, well behind GPT-5.6 Sol (max) at 32.3 and Claude Fable 5.1 (xhigh) at 31.1 percent.
Two scores actually drop against 1.2. AA-LCR falls from 83 to 79 percent. Factual accuracy in AA-Omniscience slips by up to three points, because the model more often declines to answer when it's unsure.
Neither Meta nor Artificial Analysis has named a price for the max variant yet. Larger models and an open-weights release are on the way.
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Related Articles
AIOpenAI Discloses Six New Incidents of ‘Concerning' A.I. Behavior
The artificial intelligence company also released a framework for reporting when its systems go wrong.
AIOpenAI Considers New Financing at a $1.5 Trillion Valuation
The funding round would double the value of the start-up behind ChatGPT and establish it as the world’s most valuable private company.


