Public
fmtlib / fmt
Benchmark updated: 9/4/202630 Tasks
A modern formatting library
Languages
C++95.7%Python2%CMake1.7%C0.4%Shell<0.1%Cuda<0.1%Other<0.1%
Harness | Input / Output Cost | ||||||
|---|---|---|---|---|---|---|---|
1 | 25 / 30 | $0.62 | $2/$6 | 7m46s | |||
2 | 22 / 30 | $0.49 | $1.4/$4.4 | 9m16s | |||
3 | 5 / 30 | $0.23 | $1.25/$4.25 | 2m23s |
Key Takeaways
- Grok 4.5 with Mini-SWE-agent scores 83.33% at $0.62 per test, compared with GLM 5.2 (Fireworks) at 73.33% and $0.49.
- Muse Spark 1.2 with Mini-SWE-agent resolves 5 of 30 tasks at 16.67% and $0.23 per test; this 30-task result is directional.
Cost Analysis
Cost / Test vs. Accuracy
ACCURACYCOST
Average Token Use / Test
Token Usage
InputOutputReasoningCache readCache write
GLM 5.2
Grok 4.5
Cost is the clearest tradeoff in this comparison. Grok 4.5 leads at 83.33% for $0.62 per test. Muse Spark 1.2 is the lower-cost option at 16.67% for $0.23 per test.
Latency Analysis
Latency vs. Accuracy
ACCURACYLATENCY
Average Response Time / Test
Response Time
GLM 5.2
Grok 4.5
Muse Spark 1.2
Latency separates several models with similarly strong scores. Grok 4.5 leads at 83.33%, while Muse Spark 1.2 is fastest at 2m 23s with 16.67% accuracy.