Public

elastic / curator

Benchmark updated: 9/15/202630 Tasks
Add models

Curator: Tending your Elasticsearch indices

Languages

Python99.2%Shell0.6%Dockerfile0.2%

Harness

1

Mini-SWE-agent
24 / 30

$1.03

5m38s

Key Takeaways

  • This 30-task result is directional, not a statistically significant comparison.
  • GPT-5.6 Sol with Mini-SWE-agent averages 10,198.90 output tokens and 9,128.07 reasoning tokens per task.

Cost Analysis

Cost / Test vs. Accuracy
ACCURACYCOST

Average Token Use / Test

Token Usage
InputOutputReasoningCache readCache write
GPT-5.6 Sol
999K

Cost is the clearest tradeoff in this comparison. GPT-5.6 Sol leads at 80.00% for $1.03 per test. No other model in this comparison is cheaper.

Latency Analysis

Latency vs. Accuracy
ACCURACYLATENCY

Average Response Time / Test

Response Time
GPT-5.6 Sol
5m 38s

GPT-5.6 Sol is both the most accurate and fastest model in this comparison at 80.00% and 5m 38s.