Sep 10, 2026
DeepSeek's V4.1 Flash evaluated across our benchmark suite
We evaluated DeepSeekโs DeepSeek V4.1 Flash across our benchmark suite.
-
It is the new #1 open-weight model on the Vals Index (57.86%, #15 of 56 overall), narrowly ahead of Kimi K3 (57.81%) at $0.30 per test versus $6.47, and finishing tasks in well under half the time. It is the cheapest model in the top 15.
-
It is the top open-weight model on Code Migration (45.62%, #9 of 58) at under $1 per test, while the next open-weight model, GLM 5.3 (44.22%), costs $24.91. It also takes #1 of 34 on SkillsBench (69.80% with skills, 61.66% without).
-
On Vibe Code Bench it is #2 among open-weight models (84.74%), 0.2 points behind Kimi K3 but roughly 40x cheaper ($0.41 vs $17.59 per task) and about 5x faster (15 minutes vs 1.5 hours). It is also the #2 open-weight model on Terminal-Bench 2.1 (74.53% across three full trials).
-
Compared to its predecessor, DeepSeek V4 Flash 0731, it gains 4.3 points on the Vals Index, with the largest jumps on Legal Research Bench (+11.1 points), Vibe Code Bench (+10.0) and Terminal-Bench 2.1 (+7.5). Despite list prices two to four times higher, it costs less per test on most agentic benchmarks because it finishes tasks in far fewer tokens, and it roughly halves latency on Vibe Code Bench, EMB, Legal Research and Harveyโs Legal Agent Benchmark.
The model has a 1M-token context window and supports up to 384k output tokens. Evaluations were run at temperature 1 with default top-p and high reasoning effort.
Congrats to the DeepSeek team on another strong release!