Release Date: Aug 7, 2025

Developer OpenAIΒ πŸ‡ΊπŸ‡Έ
Context Window 400k
Max Output Tokens 128k
Token Costs (in/out) $1.25/10.00
Weights Private
Input Modalities

Accuracy

63.40 %

Avg. Cost (In/Out)

$ 1.25 / $ 10.00

Latency

7 min 49 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±2.10
15/86

0.0%

Β±1.94
28/87

0.0%

Β±0.88
41/96

0.0%

Β±3.35
37/77

0.0%

Β±0.87
46/141

0.0%

Β±3.41
56/86

0.0%

Β±1.76
43/135

0.0%

Β±4.58
24/62

0.0%

Β±0.97
26/140

0.0%

Β±0.38
12/139

0.0%

Β±0.34
39/135

0.0%

Β±0.93
36/90

0.0%

Β±2.07
65/86
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : OpenAI
Temperature: Default
Top P: Default
Top K: Default
Max Output Tokens: 128,000
Reasoning Effort: high

Updates

Aug 7, 2025

We evaluated OpenAI’s newly-released GPT 5 family of models and found that GPT 5 achieves SOTA performance for a fraction of the cost compared to similarly performing models.

Of the three, GPT 5 is the strongest model in the family with SOTA performance on public LegalBench and AIME benchmarks.

On private benchmarks, GPT 5 and GPT 5 Mini achieve top 10 performance on all but CaseLaw. Most notably, GPT 5 Mini is the new SOTA model on TaxEval at a substantially lower cost.

On public benchmarks, GPT 5 places top 5 and GPT 5 Mini places top 10 on nearly everything. Further, we found the two models have complementary strengths - GPT 5 is SOTA on LegalBench and AIME, while GPT 5 Mini is SOTA on LiveCodeBench.

Lastly, GPT 5 Nano achieves middle of the pack performance across the board. It narrowly places in the top 10 on AIME, compared to GPT 5 and GPT 5 Mini which top the charts.