Anthropic's Opus 5 Sets New Benchmark with 30.2% Score
Anthropic's Claude Opus 5 scored 30.2 percent on the ARC-AGI-3 benchmark, nearly quadrupling the previous record held by GPT-5.6 Sol, which scored just 7.8 percent. The benchmark's developers noted that Opus 5 independently formulated reflection equations, a behavior not previously observed in other models.
In addition to its impressive performance, Opus 5 is cost-effective, priced at half the token rate of Fable 5 while outperforming it in most benchmarks. With a score of 61 points on the Artificial Analysis Intelligence Index, it leads in analytical quality and coding capabilities, making it an attractive option for enterprises focused on efficiency and performance.
Why it matters: With Opus 5's 30.2% score, Anthropic has positioned itself as a leader in AI performance, significantly outpacing OpenAI's GPT-5.6 Sol and offering a more cost-effective solution for enterprises looking to optimize their AI investments.
Key Takeaways
- Opus 5 costs half as much as Fable 5, making it economically attractive for users.
- It leads the Artificial Analysis Intelligence Index with a score of 61 points, surpassing both Fable 5 and GPT-5.6 Sol.
- Anthropic's advancements could compel OpenAI to enhance its models or adjust pricing strategies to maintain competitiveness.