Release Date: Apr 16, 2025

Developer OpenAI Β πŸ‡ΊπŸ‡Έ
Context Window 200k
Max Output Tokens 100k
Token Costs (in/out) $2.00/8.00
Weights Private
Input Modalities

Accuracy

72.38 %

Avg. Cost (In/Out)

$ 2.00 / $ 8.00

Latency

34.90 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±2.16
27/93

0.0%

Β±1.87
62/95

0.0%

Β±0.93
40/98

0.0%

Β±3.31
48/81

0.0%

Β±0.85
34/145

0.0%

Β±1.86
54/138

0.0%

Β±1.03
45/143

0.0%

Β±0.42
45/145

0.0%

Β±0.34
56/138

0.0%

Β±0.95
44/93
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : OpenAI
Temperature: Default
Top P: Default
Top K: Default
Max Output Tokens: 100,000
Reasoning Effort: high

Updates

Apr 18, 2025

We just evaluated o3 and o4 Mini on all benchmarks!

  • o3 achieved the #1 overall accuracy ranking on our benchmarks, with exceptional performance on complex reasoning tests like MMMU Pro (#1/22), MMLU Pro (#1/35), GPQA Diamond (#1/35) and proprietary benchmarks like TaxEval (#1/42) and CorpFin (#2/35).

  • o4 Mini achieved the second-highest accuracy across our benchmarks (82.8%), driven by strong performance on public math tests like MGSM (#1/36), MMMU Pro (#2/22), and Math500 (#4/38).

  • Legal benchmark weaknesses: Both models demonstrated significant weaknesses on our proprietary legal benchmarks, with lower ranks on ContractLaw (o3: #34/62, o4 Mini: #14/62) and CaseLaw (o3: #15/55, o4 Mini: #18/55).

  • Cost-effectiveness comparison: With similar performance levels, cost becomes a key differentiator. o4 Mini costs $4.40 for output, compared to $40.00 for o3 β€” a tenfold price difference that makes o4 Mini the more economical choice for many use cases.