Nov 24, 2025
Claude Opus 4.5 is first on Vals Index, SWE-bench Verified, SAGE.
Claude Opus 4.5 (Thinking) is the new leader on our Vals Index, surpassing its cheaper variant, Claude Sonnet 4.5 (Thinking).
Like other Anthropic models, Claude Opus 4.5 (Thinking) demonstrates strong performance on coding tasks - the thinking and nonthinking variants take the top 2 spots on SWE-bench Verified. However, the model performs worse than Claude Sonnet 4.5 (Thinking) on our new VibeCodeBench.
Unlike other Anthropic models, Claude Opus 4.5 (Thinking) showcases strong multimodal capabilities, placing first on our Multimodal Vals Index. It also places first on SAGE, surpassing the previous leader, Gemini 3 Pro (11/25). Itโs also cheaper than before - itโs a third of the price of Claude Opus 4.1 (Nonthinking) and on some benchmarks, more token efficient than Claude Sonnet 4.5 (Thinking)!
Along with Claude Sonnet 4.5 (Thinking) and Claude Opus 4.5 (Thinking), the top 3 models on our Vals Index are now all from Anthropic. We look forward to other providers catching up!