Public

avyayv / seqdataloader

Benchmark updated: 9/2/202630 Tasks
Add models

Sequence data label generation and ingestion into deep learning models

Languages

Python100%

Harness

1

Mini-SWE-agent
23 / 30

$0.47

117.73s

Key Takeaways

  • This 30-task result is directional rather than statistically significant.
  • The run averages 117.73 seconds per test, with 29,154.60 input tokens and 4,746.90 output tokens.

Cost Analysis

Cost / Test vs. Accuracy
ACCURACYCOST

Average Token Use / Test

Token Usage
InputOutputReasoningCache readCache write
openai/gpt-5.6-sol-high
258K

Cost is the clearest tradeoff in this comparison. openai/gpt-5.6-sol-high leads at 76.67% for $0.47 per test. No other model in this comparison is cheaper.

Latency Analysis

Latency vs. Accuracy
ACCURACYLATENCY

Average Response Time / Test

Response Time
openai/gpt-5.6-sol-high
1m 58s

openai/gpt-5.6-sol-high is both the most accurate and fastest model in this comparison at 76.67% and 1m 58s.