Oct 8, 2026
Anthropic's Claude Haiku 5.5 evaluated across our benchmark suite
We evaluated Anthropicโs new Claude Haiku 5.5 across our benchmark suite.
-
Haiku 5.5 places #16 of 45 on the Vals Index (54.31%) at $2.99 per test, 0.51 points behind Gemini 3.8 Flash (54.83%).
-
It takes #3 of 110 on Vibe Code Bench (90.44%), behind only Gemini 4 Argon (91.91%) and Claude Sonnet 5.5 (92.39%), and #5 of 45 on CyberBench (75.48%), ahead of Sonnet 5.5 (59.58%).
-
Further top-ten finishes on MedScribe (89.54%, #6 of 108), Vibe Code Bench 1-100 (20.41%, #7 of 22) and MysteryMechanism (35.14%, #10 of 25).
-
Compared to Claude Haiku 4.5 (Thinking), it improves from 11.39% to 90.44% on Vibe Code Bench, 7.55% to 44.57% on Code Migration, 22.66% to 59.02% on EMB and 10.58% to 43.27% on Legal Research Bench.
-
Weaker on SRE Bench (5.73%, #15 of 33), Terminal-Bench Science (10.00%), SAGE (45.78%, #38 of 91) and Harveyโs Legal Agent Benchmark (1.25%).
We ran Haiku 5.5 without a refusal fallback. Content-filter refusals are scored as failures, including 8 of 270 BioMysteryBench attempts and 1 of 70 Terminal-Bench Science tasks.
Haiku 5.5 is priced at $0.10 per million input tokens and $0.50 per million output tokens. The model has a 1M-token context window and 128k max output tokens. Evaluations were run with compute effort set to โmaxโ.
Congrats to the Anthropic team on the release!