Aug 27, 2025
Grok Code Evaluated on Coding Benchmarks!
We evaluated SpaceXAI’s Grok Code Fast on three of our coding benchmarks and found it to be much faster (and cheaper) for practical coding tasks, but significantly worse than SpaceXAI’s flagship model Grok 4 in general. Our findings are below:
- Grok Code Fast scores 62% on LCB, placing the model in the middle of the pack, comparable to other reasoning models like Claude Sonnet 4 (Nonthinking), but for a tenth of the price.
- On IOI v1, Grok Code Fast gets 4.3% placing the model at 8/12. By contrast, Grok 4 gets 26.2% and places first overall!
- On SWE-bench Verified, Grok Code Fast gets an impressive 57.6% percent, placing 4th right behind Grok 4‘s 58.6%, but with a latency of 264.68s compared to Grok 4‘s 704.8s.
Grok Code Fast is a snappier (and cheaper) model optimized for coding, and our results show that while there is significant room for improvement relative to other frontier models including SpaceXAI’s Grok 4, it performs competitively on practical coding tasks while offering benefits in terms of latency and cost.