Release Date: Mar 25, 2025

Developer GoogleΒ πŸ‡ΊπŸ‡Έ
Context Window 1M
Max Output Tokens 66k
Token Costs (in/out) $1.25/10.00
Weights Private
Input Modalities

Accuracy

78.43 %

Avg. Cost (In/Out)

$ 1.25 / $ 10.00

Latency

19.74 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±0.92
28/96

0.0%

Β±0.87
54/141

0.0%

Β±2.00
64/135

0.0%

Β±0.41
29/139

0.0%

Β±0.36
66/135

0.0%

Β±0.94
37/90
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Google
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 65,536

Updates

Mar 28, 2025

We just evaluated Gemini 2.5 Pro Exp on all benchmarks!

  • Gemini 2.5 Pro Exp is Google’s latest experimental model and the new State-of-the-Art, achieving an impressive average accuracy of 82.3% across all benchmarks with a latency of 24.68s.
  • The model ranks #1 on many of our benchmarks including CorpFin, Math500, LegalBench, GPQA Diamond, MMLU Pro, and MMMU Pro.
  • It excels in academic benchmarks, with standout performances on Math500 (95.2%), MedQA (93.0%), and MGSM (92.2%).
  • Gemini 2.5 Pro Exp demonstrates strong legal reasoning capabilities with 86.1% accuracy on CaseLaw and 83.6% on LegalBench, though it scores lower on ContractLaw (64.7%).