Release Date: Feb 19, 2026

Developer Google Β πŸ‡ΊπŸ‡Έ
Context Window 1M
Max Output Tokens 66k
Token Costs (in/out) $2.00/12.00
Weights Private
Input Modalities

Accuracy

41.90 % Β± 1.17

Cost / Test (Vals Index)

$ 1.945

Latency

9 min 20 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±1.17
37/58

0.0%

Β±3.93
38/61

0.0%

Β±4.17
26/26

0.0%

Β±2.97
37/58

0.0%

Β±1.21
41/61

0.0%

Β±2.81
39/61

0.0%

Β±2.00
2/93

0.0%

Β±1.92
65/95

0.0%

Β±0.91
6/98

0.0%

Β±4.39
22/32

0.0%

Β±3.29
24/81

0.0%

Β±0.86
58/145

0.0%

Β±2.40
19/19

0.0%

Β±4.34
54/96

0.0%

Β±1.05
1/138

0.0%

Β±0.93
6/143

0.0%

Β±0.33
3/145

0.0%

Β±0.28
4/138

0.0%

Β±0.78
10/93

0.0%

Β±0.00
32/47

0.0%

Β±1.83
29/88

0.0%

Β±1.12
21/66
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Google
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 65,536
Reasoning Effort: high

Updates

Feb 20, 2026

We evaluated Gemini 3.1 Pro Preview (02/26) across our full benchmark suite. Here are the key takeaways:

One notable metric to call out here is the model achieves this performance at a lower cost than models like Claude Opus 4.6, Claude Sonnet 4.6, GPT 5.2 and O3.

Evaluations were run with a temperature of 1.0 and a β€œhigh” thinking level, via the official Google API.

Congratulations to the Google team on another outstanding model!

Feb 19, 2026

We evaluated Gemini 3.1 Pro Preview (02/26) on our Vals Index. Here are the key takeaways:

We are evaluating Gemini 3.1 Pro Preview (02/26) on our full suite of benchmarks and will be sharing updates soon!