Public
avyayv / seqdataloader
Benchmark updated: 9/2/202630 Tasks
Sequence data label generation and ingestion into deep learning models
Languages
Python100%
Harness | ||||||
|---|---|---|---|---|---|---|
1 | 23 / 30 | $0.47 | 117.73s |
Key Takeaways
- This 30-task result is directional rather than statistically significant.
- The run averages 117.73 seconds per test, with 29,154.60 input tokens and 4,746.90 output tokens.
Cost Analysis
Cost / Test vs. Accuracy
ACCURACYCOST
Average Token Use / Test
Token Usage
InputOutputReasoningCache readCache write
openai/gpt-5.6-sol-high
Cost is the clearest tradeoff in this comparison. openai/gpt-5.6-sol-high leads at 76.67% for $0.47 per test. No other model in this comparison is cheaper.
Latency Analysis
Latency vs. Accuracy
ACCURACYLATENCY
Average Response Time / Test
Response Time
openai/gpt-5.6-sol-high
openai/gpt-5.6-sol-high is both the most accurate and fastest model in this comparison at 76.67% and 1m 58s.