Public

fmtlib / fmt

Benchmark updated: 9/4/202630 Tasks
Add models

A modern formatting library

Languages

C++95.7%Python2%CMake1.7%C0.4%Shell<0.1%Cuda<0.1%Other<0.1%

Harness

1

Mini-SWE-agent
25 / 30

$0.62

7m46s

2

Mini-SWE-agent
22 / 30

$0.49

9m16s

3

Mini-SWE-agent
5 / 30

$0.23

2m23s

Key Takeaways

  • Grok 4.5 with Mini-SWE-agent scores 83.33% at $0.62 per test, compared with GLM 5.2 (Fireworks) at 73.33% and $0.49.
  • Muse Spark 1.2 with Mini-SWE-agent resolves 5 of 30 tasks at 16.67% and $0.23 per test; this 30-task result is directional.

Cost Analysis

Cost / Test vs. Accuracy
ACCURACYCOST

Average Token Use / Test

Token Usage
InputOutputReasoningCache readCache write
GLM 5.2
2.1M
Grok 4.5
834K

Cost is the clearest tradeoff in this comparison. Grok 4.5 leads at 83.33% for $0.62 per test. Muse Spark 1.2 is the lower-cost option at 16.67% for $0.23 per test.

Latency Analysis

Latency vs. Accuracy
ACCURACYLATENCY

Average Response Time / Test

Response Time
GLM 5.2
9m 16s
Grok 4.5
7m 46s
Muse Spark 1.2
2m 23s

Latency separates several models with similarly strong scores. Grok 4.5 leads at 83.33%, while Muse Spark 1.2 is fastest at 2m 23s with 16.67% accuracy.