Independent Evaluation, Unbiased Benchmarks

Testing AI on Real-World Tasks

We benchmark the world's leading AI models on economically valuable tasks such as finance, software, and frontier risk like cybersecurity, recursive self improvement and mental health. We run all of our own evaluations and create many of our benchmarks in-house.

Sep 22, 2026
Showing best model from each labView Full Results

Latest Reports

Recent benchmark releases and model evaluations.

Sep 22, 2026

Ten Claude Opus 5.5 agents prove a faster shortest-path algorithm in Lean

Industry Leaderboard

Model performance on different sections of the economy.

Industry
Benchmark

Vibe Code Bench v1.1

Benchmark data unavailable

Benchmark data not found

Model Performance Over Time

Tracking how foundation models improve with each release

AccuracyTime
Vals IndexSep 22, 2026
Viewing 17 lab performance frontiers.
100Accuracy806040200
Oct '25Dec '25Mar '26Jun '26Aug '26