Public
elastic / curator
Benchmark updated: 9/15/202630 Tasks
Curator: Tending your Elasticsearch indices
Languages
Python99.2%Shell0.6%Dockerfile0.2%
Harness | Input / Output Cost | ||||||
|---|---|---|---|---|---|---|---|
1 | 24 / 30 | $1.03 | $4/$20 | 5m38s |
Key Takeaways
- This 30-task result is directional, not a statistically significant comparison.
- GPT-5.6 Sol with Mini-SWE-agent averages 10,198.90 output tokens and 9,128.07 reasoning tokens per task.
Cost Analysis
Cost / Test vs. Accuracy
ACCURACYCOST
Average Token Use / Test
Token Usage
InputOutputReasoningCache readCache write
GPT-5.6 Sol
Cost is the clearest tradeoff in this comparison. GPT-5.6 Sol leads at 80.00% for $1.03 per test. No other model in this comparison is cheaper.
Latency Analysis
Latency vs. Accuracy
ACCURACYLATENCY
Average Response Time / Test
Response Time
GPT-5.6 Sol
GPT-5.6 Sol is both the most accurate and fastest model in this comparison at 80.00% and 5m 38s.