Release Date: Nov 24, 2025

Developer Anthropicย ๐Ÿ‡บ๐Ÿ‡ธ
Context Window 200k
Max Output Tokens 64k
Token Costs (in/out) $5.00/25.00
Weights Private
Input Modalities

Accuracy

69.57 %

Avg. Cost (In/Out)

$ 5.00 / $ 25.00

Latency

5 min 7 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ2.01
18/91

0.0%

ยฑ1.90
21/93

0.0%

ยฑ0.92
23/98

0.0%

ยฑ3.40
7/81

0.0%

ยฑ0.85
28/145

0.0%

ยฑ3.16
62/94

0.0%

ยฑ2.48
45/138

0.0%

ยฑ1.04
47/143

0.0%

ยฑ0.39
28/143

0.0%

ยฑ0.38
31/138

0.0%

ยฑ0.90
36/93

0.0%

ยฑ1.90
38/88

0.0%

ยฑ5.31
16/67
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Anthropic
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 64,000
Compute Effort: high

Updates

Nov 24, 2025

Claude Opus 4.5 (Thinking) is the new leader on our Vals Index, surpassing its cheaper variant, Claude Sonnet 4.5 (Thinking).

Like other Anthropic models, Claude Opus 4.5 (Thinking) demonstrates strong performance on coding tasks - the thinking and nonthinking variants take the top 2 spots on SWE-bench Verified. However, the model performs worse than Claude Sonnet 4.5 (Thinking) on our new VibeCodeBench.

Unlike other Anthropic models, Claude Opus 4.5 (Thinking) showcases strong multimodal capabilities, placing first on our Multimodal Vals Index. It also places first on SAGE, surpassing the previous leader, Gemini 3 Pro (11/25). Itโ€™s also cheaper than before - itโ€™s a third of the price of Claude Opus 4.1 (Nonthinking) and on some benchmarks, more token efficient than Claude Sonnet 4.5 (Thinking)!

Along with Claude Sonnet 4.5 (Thinking) and Claude Opus 4.5 (Thinking), the top 3 models on our Vals Index are now all from Anthropic. We look forward to other providers catching up!