Feb 24, 2026
GPT 5.3 Codex Evaluated
Results are live for GPT 5.3 Codex! Overall, the model is a strong performer on programming tasks:
- On Terminal-Bench 2.0, itโs a +12.3% improvement over GPT 5.2, placing second behind Gemini 3.1 Pro Preview (02/26).
- On IOI, it performs well (at #2) but it doesnโt quite match GPT 5.2. Latency is high at nearly an hour per question, though thatโs half of GPT 5.2โs.
- On VibeCodeBench, it scores 41.4% (#4) compared to GPT 5.2โs 46.9% (#2) โ the Codex model appears less adapted to the OpenHands harness than the more general GPT 5.2.
All benchmarks were run with โxhighโ reasoning except Terminal-Bench 2.0, which used โhighโ.