Release Date: Sep 22, 2026

Developer Anthropic ย ๐Ÿ‡บ๐Ÿ‡ธ
Context Window 1M
Max Output Tokens 128k
Token Costs (in/out) $4.00/20.00
Weights Private
Input Modalities

Accuracy

66.16 % ยฑ 1.00

Cost / Test (Vals Index)

$ 19.22

Latency

82 min 57 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ1.00
4/60

0.0%

ยฑ4.61
4/63

0.0%

ยฑ2.38
2/60

0.0%

ยฑ0.17
8/63

0.0%

ยฑ3.48
4/63

0.0%

ยฑ2.27
15/95

0.0%

ยฑ1.93
1/97

0.0%

ยฑ3.36
2/13

0.0%

ยฑ0.00
1/34

0.0%

ยฑ3.36
33/83

0.0%

ยฑ1.19
3/38

0.0%

ยฑ2.92
2/11

0.0%

ยฑ3.15
6/23

0.0%

ยฑ4.83
1/20

0.0%

ยฑ1.53
2/98

0.0%

ยฑ0.98
2/13

0.0%

ยฑ4.94
2/30

0.0%

ยฑ2.75
1/48

0.0%

ยฑ1.01
1/31

0.0%

ยฑ6.02
2/27
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Anthropic
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 128,000
Compute Effort: max

Updates

Sep 22, 2026

We evaluated Anthropicโ€™s new Claude Opus 5.5 across our benchmark suite.

We ran Opus 5.5 with Claude Opus 5 and Claude Opus 4.8 as server-side fallbacks for refusals. Counting fallback-assisted tasks as failures changes Terminal-Bench 2.1 from 87.64% to 79.77% (26 of 267 tasks), Terminal-Bench 4.0 from 61.62% to 53.54% (30 of 198), Vibe Code Bench from 90.29% to 83.34% (4 of 50), Terminal-Bench Science from 48.57% to 47.14%, and MysteryMechanism from 49.55% to 49.10%. The effect is largest on SRE Bench, where 217 of 262 tasks (82.82%) were fallback-assisted and the score falls from 33.59% to 5.34%. Legal Research Bench saw 11 refusals with no fallback, which did not change its score.

Opus 5.5 is priced at $4.00 per million input tokens and $20.00 per million output tokens. The model has a 1M-token context window and 128k max output tokens. Evaluations were run with compute effort set to โ€œmaxโ€ except Terminal-Bench 2.1, which used โ€œhighโ€ effort, and temperature set to 1.0.

Congrats to the Anthropic team on the release!