Release Date: Jul 22, 2026

Developer Anthropic Β πŸ‡ΊπŸ‡Έ
Context Window 1M
Max Output Tokens 128k
Token Costs (in/out) $5.00/25.00
Weights Private
Input Modalities

Accuracy

67.21 % Β± 0.98

Cost / Test (Vals Index)

$ 18.81

Latency

55 min 49 s

Fallback Rate 0.42%
Refusal Rate 0.42%
Vals Index
BenchmarksAccuracyRankings

0.0%

Β±0.98
2/59

0.0%

Β±4.37
2/62

0.0%

3/5

0.0%

Β±2.54
24/26

0.0%

Β±2.24
3/59

0.0%

Β±0.08
7/62

0.0%

Β±0.00
5/11

0.0%

Β±3.46
2/62

0.0%

Β±1.99
1/94

0.0%

Β±1.92
2/96

0.0%

Β±0.88
1/98

0.0%

Β±3.25
3/12

0.0%

Β±0.99
3/33

0.0%

Β±3.29
19/82

0.0%

Β±2.03
4/10

0.0%

Β±2.96
2/22

0.0%

Β±0.83
23/145

0.0%

Β±4.33
1/19

0.0%

Β±3.00
4/97

0.0%

Β±2.06
2/13

0.0%

Β±1.24
8/138

0.0%

Β±9.96
5/29

0.0%

Β±0.91
4/143

0.0%

Β±0.42
7/146

0.0%

Β±0.28
2/138

0.0%

Β±0.72
2/93

0.0%

Β±1.21
3/48

0.0%

Β±4.58
8/35

0.0%

Β±0.76
1/88

0.0%

Β±2.62
3/29

0.0%

Β±5.05
3/27
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Anthropic
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 128,000
Compute Effort: max

Updates

Jul 23, 2026

We evaluated Anthropic’s new Claude Opus 5 across 27 benchmark leaderboards.

We ran Opus 5 with Claude Opus 4.8 as a server-side fallback for refusals. Counting fallback-assisted results as failures changes Terminal-Bench 2.1 from 84.64% to 81.27%, MMLU Pro from 91.59% to 91.58%, the Vals Index from 74.82% to 74.47%, and the Vals Multimodal Index from 73.90% to 73.58%. Fallbacks on Finance Agent v2 and CyberBench did not change their published scores.

The model has a 1M-token context window and 128k max output tokens. Evaluations were run with compute effort set to β€œmax” except Terminal-Bench 2.1, which used β€œhigh” effort, and temperature set to 1.0 where configurable.

Congrats to the Anthropic team on the release!