Sep 28, 2026
Anthropic's Claude Sonnet 5.5 evaluated across our benchmark suite
We evaluated Anthropicโs new Claude Sonnet 5.5 across our benchmark suite.
-
Sonnet 5.5 places #2 of 66 on the Vals Index (69.22%) at $20.80 per test, 0.47 points behind Claude Opus 5.5 (69.69%) at under two-thirds of its $32.77 per test, and ahead of Claude Fable 5.1 (68.83%) and Claude Opus 5 (67.21%).
-
It takes #1 on three leaderboards: Vibe Code Bench (92.39%, ahead of Claude Fable 5 at 90.35%), Code Migration (69.83%, ahead of GPT-6 Astra at 67.74%) and BioMysteryBench (81.11%, ahead of a three-way tie at 79.26%), and scores a perfect 100.00% on ProofBench v1.1, matching Opus 5.5, Fable 5.1 and AlephProver.
-
It sits at #3 on EMB (75.71% of 66), MedScribe (91.10% of 104), Tax Agent Bench (73.40% of 61), Terminal-Bench 4.0 (53.03% of 37), Terminal-Bench Science (38.57% of 33) and MysteryMechanism (49.10% of 19), and #4 of 14 on SRE Bench (30.15%).
-
Further top-ten finishes on Terminal-Bench 2.1 (83.15%, tied #6 of 75), IOI (83.06%, #7 of 36), Legal Research Bench (48.08%, tied #7 of 69), Finance Agent v2 (58.10%, #9 of 70) and SAGE (51.77%, #10 of 88).
-
Mid-board on MedCode (52.92%, #11 of 102) and Public Benefits Bench (67.19%, #12 of 43). Weaker on CyberBench (59.58%, #31 of 41) and Harveyโs Legal Agent Benchmark (2.92%, tied #35 of 70).
We ran Sonnet 5.5 with Claude Sonnet 5 as a server-side fallback for refusals. Counting fallback-assisted tasks as failures changes Terminal-Bench 2.1 from 83.15% to 80.52% (10 of 267 tasks), Terminal-Bench 4.0 from 53.03% to 50.51% (7 of 198) and the Vals Index from 69.22% to 68.89%. The effect is largest on security work: CyberBench falls from 59.58% to 41.97% (54 of 116 tasks), and SRE Bench from 30.15% to 19.08% (126 of 262). Refusals that the fallback also refused are scored as failures, including 41 SRE Bench tasks and nine BioMysteryBench attempts.
Sonnet 5.5 is priced at $2.00 per million input tokens and $10.00 per million output tokens. The model has a 1M-token context window and 128k max output tokens. Evaluations were run with compute effort set to โmaxโ except Terminal-Bench 2.1, which used โhighโ effort, and temperature set to 1.0.
Congrats to the Anthropic team on the release!