Aug 9, 2025
Opus 4.1 (Thinking) Evaluated!
We just evaluated Claude Opus 4.1 (Thinking) on our non-agentic benchmarks. While it placed in the top 10 on 6 of our public benchmarks, its performance on our private benchmarks was fairly mediocre.
-
On our private benchmarks, Claude Opus 4.1 (Thinking) lands squarely in the middle of the pack β barely making the top 10 on our TaxEval benchmark.
-
On public benchmarks, however, Claude Opus 4.1 (Thinking) ranks in the top 10 on 6 of the benchmarks we evaluated. Notably, it takes 2nd place on MMLU Pro behind only Claude Opus 4.1 (Nonthinking) and claims 1st place on MGSM.