Public

agentspacecoding / codeburn

Benchmark updated: 9/1/202630 Tasks
Add models

Free, local tool to track AI coding token usage and cost across 37 tools and agents (Claude Code, Cursor, Codex, Gemini and more), by model, project, and task. npx codeburn

Languages

TypeScript78.8%Swift15.8%JavaScript2.1%CSS1.6%Rust0.9%HTML0.6%Other0.2%

Harness

1

Mini-SWE-agent
9 / 30

$5.37

19m57s

Key Takeaways

  • This single supplied run establishes no model or harness comparison.

Cost Analysis

Cost / Test vs. Accuracy
ACCURACYCOST

Average Token Use / Test

Token Usage
InputOutputReasoningCache readCache write

No token usage data available.

Cost is the clearest tradeoff in this comparison. Claude Sonnet 5 leads at 30.00% for $5.37 per test. No other model in this comparison is cheaper.

Latency Analysis

Latency vs. Accuracy
ACCURACYLATENCY

Average Response Time / Test

Response Time
Claude Sonnet 5
19m 57s

Claude Sonnet 5 is both the most accurate and fastest model in this comparison at 30.00% and 19m 57s.