Oct 16, 2025
Claude Haiku 4.5 Evaluated on All Benchmarks!
We evaluated Claude Haiku 4.5 (Thinking) and found strong performance:
- The model places 3rd on our Vals Index, demonstrating well-rounded capabilities across diverse tasks.
- On Terminal-Bench, Haiku 4.5 achieves 3rd place, showing particular strength on coding tasks.
- While it performs well on certain coding benchmarks, the model achieves middle-of-the-pack performance on most other benchmarks.
- The model struggles significantly on our proprietary CaseLaw benchmark and the public MedQA, GPQA Diamond, MMLU Pro, MMMU Pro, and LiveCodeBench benchmarks.
- Compared to Claude Sonnet 4.5 (Thinking), Haiku trades some performance for significantly faster speed and lower cost - see our model comparison for details.
Overall, Haiku 4.5 sits firmly on the Pareto frontier.