Release Date: Aug 28, 2026

Developer Anthropicย ๐Ÿ‡บ๐Ÿ‡ธ
Context Window 1M
Max Output Tokens 128k
Token Costs (in/out) $10.00/50.00
Weights Private
Input Modalities

Accuracy

67.87 % ยฑ 1.10

Cost / Test (Vals Index)

$ 28.40

Latency

72 min 25 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ1.10
1/51

0.0%

ยฑ2.08
1/51

0.0%

ยฑ2.06
3/54

0.0%

ยฑ3.46
2/54

0.0%

ยฑ2.17
7/88

0.0%

ยฑ1.95
1/90

0.0%

ยฑ0.90
2/97

0.0%

ยฑ0.00
1/27

0.0%

ยฑ3.34
25/78

0.0%

ยฑ1.13
2/32

0.0%

ยฑ2.60
2/7

0.0%

ยฑ0.83
9/144

0.0%

ยฑ1.57
2/89

0.0%

ยฑ1.86
8/137

0.0%

ยฑ0.86
1/142

0.0%

ยฑ0.42
2/141

0.0%

ยฑ0.27
1/137

0.0%

ยฑ0.70
1/92

0.0%

ยฑ4.76
5/32

0.0%

ยฑ0.99
2/59
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Anthropic
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 128,000
Compute Effort: max

Updates

Sep 1, 2026

We evaluated Anthropicโ€™s new Claude Fable 5.1 across our benchmark suite.

Refusals remain frequent on bio- and cyber-adjacent tasks, so we ran Fable 5.1 with Claude Opus 5 and Claude Opus 4.8 as server-side fallbacks. Counting fallback-assisted tasks as failures changes Terminal-Bench 2.1 from 85.02% to 79.03% (23 of 267 tasks), the Vals Index from 67.87% to 66.85%, Legal Research Bench from 55.29% to 54.33%, and Harveyโ€™s Legal Agent Benchmark from 6.67% to 5.83%. The effect is largest on ReverseEngBench, where 195 of 262 tasks (74.43%) were fallback-assisted and the score falls from 22.90% to 10.69%.

The model has a 1M-token context window and 128k max output tokens. Evaluations were run with compute effort set to โ€œmaxโ€ and temperature set to 1.0.

Congrats to the Anthropic team on the release!