Independent Evaluation, Unbiased Benchmarks

Testing AI on Real-World Tasks

We benchmark the world's leading AI models on rigorous, domain-specific tasks in finance, law, software, healthcare, and more. We run all of our own evaluations and create many of our benchmarks in-house.

Vals AI Updates

Fresh updates from our testing queue

model
07/21/2026

Google's Gemini 3.6 Flash evaluated on the Vals Index

Google's Gemini 3.6 Flash evaluated on the Vals Index

View Details

Benchmarks

Accuracy

Rankings

62.43%

± 1.84
14/ 38

25.00%

± 3.01
18/ 25

53.15%

± 2.16
7/ 74

79.66%

± 1.86
33/ 73

48.69%

± 3.38
19/ 69

3.33%

± 1.17
12/ 25

73.78%

± 1.63
8/ 43
Contact us
Or send us an email at contact@vals.ai

License type:

Proprietary (contact us to get access)
Industry Partner
Academic

Read our methodology.

Industry Leaderboard

Independent benchmarks for industry-specific AI performance.

Industry
Benchmark

Model Performance Over Time

Tracking how foundation models improve with each release

85%73%61%49%37%25%
Oct '25Nov '25Jan '26Mar '26Apr '26Jun '26Aug '26