Independent Evaluation, Unbiased Benchmarks

Testing AI on Real-World Tasks

We benchmark the world's leading AI models on economically valuable tasks such as finance, software, and frontier risk like cybersecurity, recursive self improvement and mental health. We run all of our own evaluations and create many of our benchmarks in-house.

Oct 02, 2026
Showing best model from each labView Full Results

Latest Reports

Recent benchmark releases and model evaluations.

Sep 30, 2026

Has the Bitter Lesson Come for AI Detectors?

AI Detection BenchmarkVals
SystemsAccuracy
1Claude Opus 5.5
98.4% ±0.62
2GPT-6 Astra
95.7% ±0.92
3Originality.ai
94.1% ±1.18
4Pangram 4
93.9% ±1.21
5GPTZero
78.1% ±1.82
6Sapling
77.5% ±1.63
7DeepSeek V4.1 Flash
65.4% ±1.72
8Copyleaks
64.6% ±1.44

Industry Leaderboard

Model performance on different sections of the economy.

Industry
Benchmark

Vibe Code Bench v1.1

Benchmark data unavailable

Benchmark data not found

Model Performance Over Time

Tracking how foundation models improve with each release

AccuracyTime
Vals Index•Oct 02, 2026
Viewing 16 lab performance frontiers.
100Accuracy806040200
Feb '26Apr '26May '26Jul '26Sep '26