Public

dwojtaszek / zipper

Benchmark updated: 9/15/202630 Tasks
Add models

Generate industry standard load files

Languages

C#76.3%Shell11.6%Batchfile6.6%Python5.5%

Harness

1

Mini-SWE-agent
28 / 30

$1.55

30m43s

2

Mini-SWE-agent
26 / 30

$4.54

18m34s

3

Mini-SWE-agent
13 / 30

$0.38

5m33s

Key Takeaways

  • Claude Opus 5 with Mini-SWE-agent resolves 26 of 30 tasks, while Claude Haiku 4.5 (Thinking) resolves 13 of 30.
  • Claude Haiku 4.5 (Thinking) with Mini-SWE-agent has the lowest latency at 332.59 seconds; GPT-5.6 Sol costs $1.55 per test versus $4.54 for Claude Opus 5.

Cost Analysis

Cost / Test vs. Accuracy
ACCURACYCOST

Average Token Use / Test

Token Usage
InputOutputReasoningCache readCache write

No token usage data available.

Cost is the clearest tradeoff in this comparison. GPT-5.6 Sol leads at 93.33% for $1.55 per test. Claude Haiku 4.5 (Thinking) is the lower-cost option at 43.33% for $0.38 per test.

Latency Analysis

Latency vs. Accuracy
ACCURACYLATENCY

Average Response Time / Test

Response Time
GPT-5.6 Sol
30m 43s
Claude Opus 5
18m 34s
Claude Haiku 4.5 (Thinking)
5m 33s

Latency separates several models with similarly strong scores. GPT-5.6 Sol leads at 93.33%, while Claude Haiku 4.5 (Thinking) is fastest at 5m 33s with 43.33% accuracy.