Public

punt-labs / lux

Benchmark updated: 9/14/202630 Tasks
Add models

GUI - a whiteboard and dashboard interface for agents to share visual data with humans.

Languages

Python98.3%Shell0.8%TeX0.8%Makefile0.1%

Harness

1

Mini-SWE-agent
22 / 30

$1.19

7m36s

2

Mini-SWE-agent
22 / 30

$2.20

20m07s

3

Mini-SWE-agent
18 / 30

$0.82

11m23s

Key Takeaways

  • This 30-task result is directional: Claude Opus 5 (High Effort) and Muse Spark 1.2 each score 73.33%, while GLM 5.2 (Fireworks) scores 60%.
  • Muse Spark 1.2 with Mini-SWE-agent has lower cost per test and latency than Claude Opus 5 (High Effort): $1.19 and 456 seconds versus $2.20 and 1207 seconds.
  • GLM 5.2 (Fireworks) with Mini-SWE-agent is the lowest-cost run at $0.82 per test, but resolves four fewer tasks than the two leading runs.

Cost Analysis

Cost / Test vs. Accuracy
ACCURACYCOST

Average Token Use / Test

Token Usage
InputOutputReasoningCache readCache write
Muse Spark 1.2
5.3M
GLM 5.2
3.7M
anthropic/claude-opus-5-high
2.4M

Cost is the clearest tradeoff in this comparison. anthropic/claude-opus-5-high leads at 73.33% for $2.20 per test. Muse Spark 1.2 is the lower-cost option at 73.33% for $1.19 per test.

Latency Analysis

Latency vs. Accuracy
ACCURACYLATENCY

Average Response Time / Test

Response Time
anthropic/claude-opus-5-high
20m 7s
GLM 5.2
11m 23s
Muse Spark 1.2
7m 36s

Latency separates several models with similarly strong scores. anthropic/claude-opus-5-high leads at 73.33%, while Muse Spark 1.2 is fastest at 7m 36s with 73.33% accuracy.