Release Date: Jul 9, 2026

Developer Metaย ๐Ÿ‡บ๐Ÿ‡ธ
Context Window 1M
Max Output Tokens 256k
Token Costs (in/out) $1.25/4.25
Weights Private
Input Modalities

Accuracy

54.75 % ยฑ 1.15

Cost / Test (Vals Index)

$ 1.251

Latency

13 min 48 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ1.15
20/54

0.0%

ยฑ4.07
22/56

0.0%

ยฑ3.15
27/54

0.0%

ยฑ0.79
9/57

0.0%

ยฑ0.00
8/8

0.0%

ยฑ3.37
21/57

0.0%

ยฑ1.95
5/92

0.0%

ยฑ0.93
37/98

0.0%

ยฑ3.28
32/80

0.0%

ยฑ1.27
18/33

0.0%

ยฑ0.77
2/145

0.0%

ยฑ3.33
19/92

0.0%

ยฑ2.10
23/138

0.0%

ยฑ0.99
29/143

0.0%

ยฑ0.42
24/142

0.0%

ยฑ0.32
17/138

0.0%

ยฑ0.82
18/93

0.0%

ยฑ0.00
27/45

0.0%

ยฑ4.44
9/33

0.0%

ยฑ1.72
21/88

0.0%

ยฑ0.99
22/62
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Meta
Temperature: Default
Top P: Default
Top K: Default
Max Output Tokens: 256,000
Reasoning Effort: xhigh

Updates

Jul 9, 2026

We evaluated Metaโ€™s new Muse Spark 1.1 across our benchmark suite.

  • It debuts at #14 on the Vals Index (54.75%), narrowly behind GPT 5.5 (57.41%). At $1.25/test, it is the fourth-lowest-cost model in the top 20; it also averaged 827.8 seconds on the index, roughly 2โ€“4ร— faster than the top three models.

  • It takes #2 on Finance Agent v2 (57.21%), just 0.65 points behind Gemini 3.5 Flash.

  • The largest generational gain is on Vibe Code Bench. Compared with Muse Spark, Muse Spark 1.1 rises from #43 to #5 and improves by 52.48 points (19.67% โ†’ 72.16%).

  • Muse Spark 1.1 sets new highs on domain-specific work: #1 on MedScribe (88.89%), #1 on TaxEval v2 (79.72%), and #1 on Harveyโ€™s Legal Agent Benchmark (20.00%). On MedScribe, it is nearly twice as fast as the #2 model.

The model has a 1M-token context window and supports up to 256k output tokens. Evaluations were run with reasoning effort set to โ€œxhighโ€ and default temperature and top-p.

Congrats to the Meta team on the release!