Mar 5, 2026
GPT 5.4 evaluated on our full benchmark suite
We evaluated GPT 5.4 (xhigh) across our full benchmark suite.
- GPT 5.4 (xhigh) ranks #4 on both our Vals Index and #3 on the Vals Multimodal Index.
- It takes #1 on Vibe Code Bench, Proof Bench, and IOI.
- On Proof Bench, GPT 5.4 (xhigh) reaches 56% accuracy (up from sub 20% for prior OpenAI models, and beating Claude Opus 4.6 (Thinking) by 6%).
- Relative to GPT 5.2, we see especially large gains on coding/math-heavy benchmarks, including Vibe Code Bench (+13.9%), IOI (+13.0%), and Terminal-Bench 2.0 (+6.7%), along with a smaller margin of improvement on SWE-bench Verified (+1.8%).
- We also observe a few areas where GPT 5.4 (xhigh) is roughly flat or slightly behind GPT 5.2 in our current data (for example, GPQA Diamond is tied, while LCB and Finance Agent have GPT 5.2 slightly higher).
Evaluations were run via the official OpenAI API using โxhighโ reasoning, except on Terminal-Bench 2.0, where we used โhighโ.