Release Date: Mar 3, 2026

Developer Google ย ๐Ÿ‡บ๐Ÿ‡ธ
Context Window 1M
Max Output Tokens 66k
Token Costs (in/out) $0.25/1.50
Weights Private
Input Modalities

Accuracy

15.46 % ยฑ 0.42

Cost / Test (Vals Index)

$ 0.136

Latency

2 min 43 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ0.42
56/58

0.0%

ยฑ1.07
57/61

0.0%

ยฑ1.09
58/58

0.0%

ยฑ0.64
55/61

0.0%

ยฑ1.25
57/61

0.0%

ยฑ2.07
25/93

0.0%

ยฑ1.82
89/95

0.0%

ยฑ0.91
20/98

0.0%

ยฑ3.48
18/81

0.0%

ยฑ0.88
74/145

0.0%

ยฑ0.00
93/96

0.0%

ยฑ1.97
66/138

0.0%

ยฑ1.08
71/143

0.0%

ยฑ0.41
44/145

0.0%

ยฑ0.34
45/138

0.0%

ยฑ0.91
37/93

0.0%

ยฑ0.00
46/47

0.0%

ยฑ2.16
75/88

0.0%

ยฑ0.99
62/66
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Google
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 65,536
Reasoning Effort: high

Updates

Mar 3, 2026

We evaluated Gemini 3.1 Flash Lite Preview across our full benchmark suite. The model is Googleโ€™s fast and cost-efficient offering.

  • The model does well on two of our multimodal benchmarks, SAGE (ranked 5th, with 49.5% accuracy) and Mortgage Tax (ranked 7th, 67.8% accuracy).
  • It ranks 14th on MMLU Pro with 86.2% accuracy.
  • On coding tasks the model still has room for improvementโ€”it ranks 29th on both Live Code Bench and Terminal-Bench 2.0, and 28th on SWE-bench Verified. It scores 0% on Vibe Code Bench.
  • Overall, it places 15th/20 on the Vals Multimodal Index and 22nd/31 on the Vals Index.
  • The cost savings compared to other models in the Gemini 3 series or Gemini 2.5 are dramatic: roughly 5โ€“20x cheaper per test across benchmarks, while maintaining respectable accuracy. For example, on Finance Agent it costs \$0.072 per test vs \$0.370 for Gemini 3 Flash, while performing comparably.

Our results show that the model does not perform as well as other models in the Gemini 3 series. However, it is fast and quite cost-efficient relative to those models, making it a good choice for applications that demand scale, speed, or cost-efficiency.

Evaluations were run with a temperature of 1.0 and a โ€œhighโ€ thinking level, via the official Google API.