Release Date: Jul 21, 2026

Developer Googleย ๐Ÿ‡บ๐Ÿ‡ธ
Context Window 1M
Max Output Tokens 66k
Token Costs (in/out) $1.50/7.50
Weights Private
Input Modalities

Accuracy

55.35 % ยฑ 1.09

Cost / Test (Vals Index)

$ 3.030

Latency

22 min 8 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ1.09
19/54

0.0%

ยฑ4.07
23/56

0.0%

ยฑ4.51
19/24

0.0%

ยฑ2.65
13/54

0.0%

ยฑ0.18
11/57

0.0%

ยฑ3.01
34/57

0.0%

ยฑ2.16
10/90

0.0%

ยฑ1.86
48/92

0.0%

ยฑ0.92
24/98

0.0%

ยฑ3.38
23/80

0.0%

ยฑ1.29
25/33

0.0%

ยฑ0.85
26/145

0.0%

ยฑ4.29
28/92

0.0%

ยฑ1.61
10/10

0.0%

ยฑ1.33
7/138

0.0%

ยฑ0.94
8/143

0.0%

ยฑ0.41
10/142

0.0%

ยฑ0.30
13/138

0.0%

ยฑ0.77
7/93

0.0%

ยฑ0.00
26/45

0.0%

ยฑ1.80
26/88

0.0%

ยฑ1.63
15/62
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Google
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 65,536
Reasoning Effort: high

Updates

Jul 21, 2026

We evaluated Googleโ€™s Gemini 3.6 Flash on the Vals Index and proprietary benchmarks.

  • Gemini 3.6 Flash scores 55.35% on the Vals Index, placing #13 of 43 models overall and finishing within 2.28 points of Gemini 3.5 Flash.

  • Gemini 3.6 Flash ranks #1 on the CyberBench Patch track at 84.75%, which measures fixing real-world open-source vulnerabilities.

  • Its coding results include 77.45% on the SWE-bench Verified Vals Index subset and 57.88% on the Vibe Code Bench Vals Index subset, improvements of 1.96 and 3.17 points over Gemini 3.5 Flash, respectively.

  • Gemini 3.6 Flash scores 73.78% across three full trials of Terminal-Bench 2.1, placing #8 of 43 models. It also places #7 of 74 models on MedCode, scoring 53.15%.

We evaluated Gemini 3.6 Flash with temperature 1, high reasoning effort, and up to 65k output tokens. The model supports a 1M-token context window, multimodal inputs, and tool calling.

Congrats to the Google team on the release!