Public
dwojtaszek / zipper
Benchmark updated: 9/15/202630 Tasks
Generate industry standard load files
Languages
C#76.3%Shell11.6%Batchfile6.6%Python5.5%
Harness | Input / Output Cost | ||||||
|---|---|---|---|---|---|---|---|
1 | 28 / 30 | $1.55 | $4/$20 | 30m43s | |||
2 | 26 / 30 | $4.54 | $5/$25 | 18m34s | |||
3 | 13 / 30 | $0.38 | $1/$5 | 5m33s |
Key Takeaways
- Claude Opus 5 with Mini-SWE-agent resolves 26 of 30 tasks, while Claude Haiku 4.5 (Thinking) resolves 13 of 30.
- Claude Haiku 4.5 (Thinking) with Mini-SWE-agent has the lowest latency at 332.59 seconds; GPT-5.6 Sol costs $1.55 per test versus $4.54 for Claude Opus 5.
Cost Analysis
Cost / Test vs. Accuracy
ACCURACYCOST
Average Token Use / Test
Token Usage
InputOutputReasoningCache readCache write
No token usage data available.
Cost is the clearest tradeoff in this comparison. GPT-5.6 Sol leads at 93.33% for $1.55 per test. Claude Haiku 4.5 (Thinking) is the lower-cost option at 43.33% for $0.38 per test.
Latency Analysis
Latency vs. Accuracy
ACCURACYLATENCY
Average Response Time / Test
Response Time
GPT-5.6 Sol
Claude Opus 5
Claude Haiku 4.5 (Thinking)
Latency separates several models with similarly strong scores. GPT-5.6 Sol leads at 93.33%, while Claude Haiku 4.5 (Thinking) is fastest at 5m 33s with 43.33% accuracy.