Sep 22, 2026
Anthropic's Claude Opus 5.5 evaluated across our benchmark suite
We evaluated Anthropicโs new Claude Opus 5.5 across our benchmark suite.
-
Opus 5.5 places #4 of 60 on the Vals Index (66.16%) at $19.22 per test, behind Claude Fable 5.1 (68.83%), Claude Opus 5 (67.21%) and GPT-6 Astra (66.61%).
-
It takes #1 on six leaderboards: Terminal-Bench 2.1 (87.64%, ahead of GPT-6 Astra at 87.27%), Terminal-Bench 4.0 (61.62%), ProgramBench (18.50% fully resolved), MedScribe (91.43%), the Vals RSI Index (37.13%), and a perfect 100.00% on ProofBench v1.1, matching Fable 5.1 and the specialized prover AlephProver.
-
It sits at #2 on Vibe Code Bench (90.29%, within noise of Claude Fable 5โs leading 90.35%), IOI (95.06%), EMB (75.94%, behind Fable 5.1 at 76.67%), MysteryMechanism (49.55%), Terminal-Bench Science (48.57%) and SRE Bench (33.59%), and matches GPT-6 Astraโs leading 79.26% on BioMysteryBench.
-
Mid-board on Public Benefits Bench (70.64%, #3 of 38), Legal Research Bench (50.48%, #4 of 63), Tax Agent Bench (70.50%, #6 of 23) and Finance Agent v2 (58.59%, #8 of 63).
-
Agentic legal work remains the weak spot: 3.75% on Harveyโs Legal Agent Benchmark (#31 of 64). It also lands below Opus 5 on MedCode (49.80%, #15 of 95) and SAGE (45.83%, #33 of 83).
We ran Opus 5.5 with Claude Opus 5 and Claude Opus 4.8 as server-side fallbacks for refusals. Counting fallback-assisted tasks as failures changes Terminal-Bench 2.1 from 87.64% to 79.77% (26 of 267 tasks), Terminal-Bench 4.0 from 61.62% to 53.54% (30 of 198), Vibe Code Bench from 90.29% to 83.34% (4 of 50), Terminal-Bench Science from 48.57% to 47.14%, and MysteryMechanism from 49.55% to 49.10%. The effect is largest on SRE Bench, where 217 of 262 tasks (82.82%) were fallback-assisted and the score falls from 33.59% to 5.34%. Legal Research Bench saw 11 refusals with no fallback, which did not change its score.
Opus 5.5 is priced at $4.00 per million input tokens and $20.00 per million output tokens. The model has a 1M-token context window and 128k max output tokens. Evaluations were run with compute effort set to โmaxโ except Terminal-Bench 2.1, which used โhighโ effort, and temperature set to 1.0.
Congrats to the Anthropic team on the release!