Anthropic's Claude Opus 5 Scores 30.2% on ARC-AGI-3 Benchmark
Anthropic's Claude Opus 5 scored 30.2 percent on the ARC-AGI-3 benchmark, nearly quadrupling the previous record of 7.8 percent held by GPT-5.6 Sol. The benchmark's developers noted that Opus 5 independently formulated reflection equations, a behavior previously unseen in AI models.
This performance leap positions Anthropic as a strong competitor in AI intelligence, challenging rivals like OpenAI and Google to enhance their own models to keep pace.
Why it matters: With Claude Opus 5's score of 30.2%, Anthropic has significantly raised the bar for AI intelligence, forcing OpenAI and Google to accelerate their research and development to remain competitive.
Key Takeaways
- Opus 5's score of 30.2% nearly quadruples the previous record of 7.8% set by GPT-5.6 Sol.
- The model's ability to independently formulate reflection equations marks a breakthrough in AI capabilities.
- OpenAI and Google must respond quickly to this advancement to maintain their market positions.