Meta sells the hosted Muse Spark 1.2 endpoint in one register — coding — but when Vals AI ran it through a common harness alongside 44 other models, it came first of 44 on Finance Agent (v2), first of 136 on TaxEval v2 and first of 31 on Harvey’s Legal Agent Benchmark, while placing fourteenth of 50 on the very Terminal-Bench the launch led with. That is the most interesting unremarked fact about this release, and it changes who should be looking at this model. For the model-versus-agent framing that explains part of the gap, see our agent-vs-model comparison.
The launch materials leave no doubt about the intended positioning: multi-file refactors, whole-repository generation, extended debugging, a terminal agent shipped alongside it, a Terminal-Bench number in the headline. The independent evidence points somewhere else.
The full independent picture
Vals AI’s distinguishing property is that every model runs through the same harness. That removes the variable that makes vendor benchmark decks hard to interpret — the fact that each system is usually tested inside its own scaffolding.
Under those conditions, Muse Spark 1.2 places 5th of 45 on the Vals Index at 71.88% ± 1.12, at $0.69 per test — the cheapest per-test cost of anything in the top five. That’s a genuinely strong headline. The domain breakdown is where it gets strange:
• Finance Agent (v2) — #1 of 44
• TaxEval v2 — #1 of 136
• Harvey’s Legal Agent Benchmark — #1 of 31
• MedScribe — #2 of 80
• CorpFin v2 — #5 of 131
• Vals Multimodal Index — #6 of 32
• SWE-bench — #9 of 79
• Terminal-Bench 2.1 — #14 of 50
• MMLU Pro — #15 of 129
Three firsts, and none of them in software engineering.
Why this probably happens
Two explanations, and they’re not mutually exclusive.
The harness explains part of it. Muse Spark 1.2 was co-trained with Muse Code, Meta’s terminal agent, on trajectories sampled from that agent’s own harness plus long-horizon repository tasks. A model trained inside a specific scaffold performs best inside that scaffold. Run it in a neutral harness and some of the advantage evaporates — which is precisely what a fourteenth-place Terminal-Bench finish looks like when the vendor claims 82.9% in its own setup.
There’s supporting arithmetic. Meta claims a 6.7-point generational gain to reach 82.9%; 82.9 − 6.7 = 76.2, exactly the official Terminal-Bench leaderboard row for the *previous* model under Princeton’s deliberately minimal mini-SWE-agent scaffold. Meta never published its baseline harness, so this is an inference. But it suggests the claimed gain may fold a harness upgrade in with the model upgrade.
The capability profile explains the rest. Look at what finance-agent, tax and legal-agent benchmarks actually require: read a large volume of documents, hold a great deal of context, follow rules precisely, reason across many steps, and produce a structured, defensible answer. That is a long-horizon, high-context, rule-following profile — and it happens to be exactly what Meta trained for when it optimised the model for whole-repository work and multi-step tool use.
The training target was coding. The capability that emerged generalises well beyond it, and arguably lands harder somewhere else.

So is it bad at coding? No.
Being precise matters here, because “fourteenth on Terminal-Bench” is easy to misread.
#9 of 79 on SWE-bench is a strong result. Fourteenth of 50 on Terminal-Bench, in a field containing every frontier model, is respectable. The model is a capable coding model. What it isn’t is the *leading* coding model, which is what a 82.9% vendor figure quoted without context implies.
The corroborating evidence points the same way. On the official Terminal-Bench 2.1 leaderboard there is no Muse Spark 1.2 entry at all — the top row is Claude Code with Claude Fable 5 at 83.8%, and the only Muse row is version 1.1 at 76.2%. Artificial Analysis scores the model at 57 on its Intelligence Index, #12 of 185, with the largest generational movement in *agentic* evaluations: a 260-point Elo jump on GDPval-AA v2 to 1631, ranking 5th.
Agentic. Not coding specifically. Every independent source keeps pointing at the same thing.
What this means for your shortlist
If you’re evaluating coding assistants, keep it on the list but stop treating it as the presumptive leader. Test it against models with audited Terminal-Bench and SWE-bench results, in *your* harness, on *your* repository. The vendor number was produced in a scaffold you don’t have.
If you’re building document-heavy professional workflows — financial analysis, tax computation, contract review, clinical documentation — this model deserves to be on a shortlist it probably isn’t on. It is independently ranked first in three of those domains, at the lowest per-test cost in the Vals top five, with a 1,048,576-token context window that suits document-set work almost perfectly.
Either way, budget for latency. OrcaRouter’s seven-day production telemetry shows a p50 first token of 7.73 seconds and p95 of 10.00 seconds; Vals’ independent runs averaged around 610 seconds per test. Artificial Analysis publishes no speed figure for the model at all. This is background-job infrastructure, not interactive infrastructure — which happens to fit document-analysis pipelines rather well and interactive coding rather poorly.
The cheap way to test any of this is to put the model behind the same interface as your current one and run a real comparison. On OrcaRouter it sits alongside 200-plus other models on one OpenAI-compatible key at 0% markup, with automatic failover, so a domain evaluation costs an afternoon rather than a procurement cycle.

The takeaway
Muse Spark 1.2 was built as a coding model and marketed as one, and it is a solid coding model — ninth of 79 on SWE-bench under a neutral harness. But the only independent multi-domain evaluation available ranks it first in finance, tax and legal agent work while placing it fourteenth at coding, and every other independent source shows its biggest gains in agentic rather than code-specific evaluations. If you were going to skip it because you already have a coding model you like, that may be the wrong reason to skip it. And if you work in a document-heavy regulated domain, it is currently the best-ranked model on three benchmarks you’ve probably never seen a model marketed against.
Sourcing note: the 82.9% Terminal-Bench figure is Meta’s own vendor-run result, unreproduced by third parties. All domain ranks, the Vals Index score and per-test cost are from Vals AI’s common-harness evaluations. Index score and agentic sub-results are from Artificial Analysis; the 76.2% and 83.8% rows from the official Terminal-Bench 2.1 leaderboard; first-token latency is OrcaRouter’s own seven-day production telemetry. The 82.9 − 6.7 = 76.2 observation is our inference. Checked August 7, 2026.
Read more: Why Companies That Scale Successfully Invest in Business Analysis Before Writing Code
Peace Of Mind Is Built, Not Bought
What Do You Have to Know About the “No Wagering” Bonuses and What Are Their Catches












