Release Date: Sep 23, 2025

Developer OpenAIΒ πŸ‡ΊπŸ‡Έ
Context Window 400k
Max Output Tokens 128k
Token Costs (in/out) $1.25/10.00
Weights Private
Input Modalities

Accuracy

84.72 %

Avg. Cost (In/Out)

$ 1.25 / $ 10.00

Latency

2 min 15 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±1.03
37/143
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : OpenAI
Temperature: Default
Top P: Default
Top K: Default
Max Output Tokens: 128,000
Reasoning Effort: high

Updates

Sep 24, 2025

We evaluated GPT 5 Codex across Terminal-Bench, SWE-bench Verified, IOI v1, and LCB, finding the following:

  • On Terminal-Bench, GPT-5 Codex takes 1st place with 58.8% accuracy, a 10% improvement over the previous #1, GPT-5. It also delivers lower cost and latency compared to the other top three models.
  • GPT 5 Codex also gets first place on SWE-bench Verified, narrowly outperforming GPT 5, which ranks 2nd by less than a percentage point.
  • GPT 5 Codex is the first model we’ve seen receive full credit on a single question on IOI v1, though overall accuracy remains low (9.8%).
  • GPT 5 Codex places second on LCB, behind GPT 5 Mini.

GPT 5 Codex is optimized for agentic coding, particularly within OpenAI’s Codex offering. For standardization, we used the same prompts and templates as we used with other models when running our evaluations.