Release Date: Sep 29, 2025

Developer Anthropicย ๐Ÿ‡บ๐Ÿ‡ธ
Context Window 1M
Max Output Tokens 64k
Token Costs (in/out) $3.00/15.00
Weights Private
Input Modalities

Accuracy

61.39 %

Avg. Cost (In/Out)

$ 3.00 / $ 15.00

Latency

9 min 43 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ2.00
32/90

0.0%

ยฑ1.87
29/92

0.0%

ยฑ0.95
50/98

0.0%

ยฑ3.21
56/80

0.0%

ยฑ0.86
51/145

0.0%

ยฑ3.73
58/92

0.0%

ยฑ2.25
63/138

0.0%

ยฑ5.92
25/62

0.0%

ยฑ1.13
83/143

0.0%

ยฑ0.45
35/142

0.0%

ยฑ0.39
28/138

0.0%

ยฑ0.97
48/93

0.0%

ยฑ2.05
62/88
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Anthropic
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 64,000

Updates

Sep 29, 2025

We ran the recently-released Claude Sonnet 4.5 (Thinking) our benchmarks, and found very strong performance:

  • On Finance Agent, it beats the previous state-of-the-art by five percentage points.
  • It also takes the #1 spot on SWE-bench Verified and Terminal-Bench, beating out GPT 5 Codex.
  • It is in the top 10 models on the majority of our benchmarks, and also showed better performance than Claude Sonnet 4 (Thinking) on almost all benchmarks.
  • It has a 1-million token context window, when the flag โ€œcontext-1m-2025-08-07โ€ is enabled

Overall, this model is extremely capable, at the same mid-range price point as its predecessor.