Release Date: Sep 28, 2026

Developer Anthropic ย ๐Ÿ‡บ๐Ÿ‡ธ
Context Window 1M
Max Output Tokens 128k
Token Costs (in/out) $2.00/10.00
Weights Private
Input Modalities

Accuracy

69.22 % ยฑ 0.96

Cost / Test (Vals Index)

$ 20.80

Latency

70 min 7 s

Fallback Rate 2.27%
Refusal Rate 0.14%
Vals Index
BenchmarksAccuracyRankings

69.22%

ยฑ0.96
2/66

69.83%

ยฑ4.26
1/69

59.58%

ยฑ5.21
31/41

75.71%

ยฑ2.44
3/66

58.10%

ยฑ0.67
9/70

48.08%

ยฑ3.47
9/69

52.92%

ยฑ2.12
11/102

91.10%

ยฑ1.96
3/104

49.10%

ยฑ3.36
3/19

100.00%

ยฑ0.00
1/41

51.77%

ยฑ3.41
10/88

67.19%

ยฑ1.22
12/43

30.15%

ยฑ2.84
4/14

73.39%

ยฑ2.98
3/61

92.39%

ยฑ1.26
1/104

81.11%

ยฑ0.64
1/19

2.92%

ยฑ1.36
36/70

83.06%

ยฑ3.74
7/36

53.03%

ยฑ1.51
3/37

38.57%

ยฑ5.86
3/33
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Anthropic
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 128,000
Compute Effort: max

Compared with

Updates

Sep 28, 2026

We evaluated Anthropicโ€™s new Claude Sonnet 5.5 across our benchmark suite.

We ran Sonnet 5.5 with Claude Sonnet 5 as a server-side fallback for refusals. Counting fallback-assisted tasks as failures changes Terminal-Bench 2.1 from 83.15% to 80.52% (10 of 267 tasks), Terminal-Bench 4.0 from 53.03% to 50.51% (7 of 198) and the Vals Index from 69.22% to 68.89%. The effect is largest on security work: CyberBench falls from 59.58% to 41.97% (54 of 116 tasks), and SRE Bench from 30.15% to 19.08% (126 of 262). Refusals that the fallback also refused are scored as failures, including 41 SRE Bench tasks and nine BioMysteryBench attempts.

Sonnet 5.5 is priced at $2.00 per million input tokens and $10.00 per million output tokens. The model has a 1M-token context window and 128k max output tokens. Evaluations were run with compute effort set to โ€œmaxโ€ except Terminal-Bench 2.1, which used โ€œhighโ€ effort, and temperature set to 1.0.

Congrats to the Anthropic team on the release!