The Public Standard for
Real World AI Performance
Generic benchmarks only go so far.
Vals AI evaluates models on the
real tasks each industry relies on.
Vals Index
Vals Multimodal Index
Benchmark consisting of a weighted performance across finance, coding, and education tasks. Showing the potential impact that LLMs can have on the economy.
Legal
CaseLaw v2
Private question-answer benchmark over Canadian court-cases.
Finance
CorpFin v2
A private benchmark evaluating understanding of long-context credit agreements
Healthcare
MedQA
Evaluating language model bias in medical questions.
Math
AIME
Challenging national math exam given to top high-school students
MATH 500
Academic math benchmark on probability, algebra, and trigonometry
MGSM
A multilingual benchmark for mathematical questions.
Academic
Education
Coding
Terminal-Bench 2.0
State-of-the-art set of difficult terminal-based tasks
Social Mobility
Public Benefits Bench v1.1
Can AI help people navigate SNAP benefits?
Top Models
Claude Opus 5
Claude Fable 5
Muse Spark 1.2
Public Benefits Bench v1
Can AI help people navigate SNAP benefits?