Sep 29, 2025
Sonnet 4.5 sets new SOTAs
We ran the recently-released Claude Sonnet 4.5 (Thinking) our benchmarks, and found very strong performance:
- On Finance Agent, it beats the previous state-of-the-art by five percentage points.
- It also takes the #1 spot on SWE-bench Verified and Terminal-Bench, beating out GPT 5 Codex.
- It is in the top 10 models on the majority of our benchmarks, and also showed better performance than Claude Sonnet 4 (Thinking) on almost all benchmarks.
- It has a 1-million token context window, when the flag โcontext-1m-2025-08-07โ is enabled
Overall, this model is extremely capable, at the same mid-range price point as its predecessor.