Claude Opus 5 Performance and Cost Analysis
Better Stackgo watch the original →
the gist
Claude Opus 5 offers performance competitive with Claude Fable 5 at a lower price point, achieving a 30.2% score on the ARC-AGI benchmark and showing improved UI generation and coding capabilities in local tests.
Model Performance and Benchmarks
Anthropic's Claude Opus 5 positions itself as a high-intelligence model comparable to Claude Fable 5. It achieved a 30.2% score on the ARC-AGI benchmark, a significant increase from the 1.5% score of Opus 4.8. This improvement is attributed to the model's ability to convert test problems into algebraic notation for logical reasoning. In third-party evaluations by Artificial Analysis, Opus 5 leads in the Agentic Index and performs strongly in coding tasks, trailing GPT-5.6 Soul by only 0.3 points while outperforming Fable 5 by 1.5 points.
Practical Application and Cost
In local testing involving the creation of a 3D Formula 1 game using Three.js and a full-stack finance dashboard, Opus 5 demonstrated superior UI generation compared to Fable 5 and GPT-5.6 Soul. For the finance dashboard, Opus 5 utilized a robust stack including React, React Router, Node SQLite, and Express. While artificial benchmarks suggest Opus 5 is significantly cheaper than Fable 5, local testing showed variable results, with Opus 5 occasionally incurring higher costs and longer generation times depending on token usage. The model maintains a pricing structure of $5 per million input tokens and $25 per million output tokens.
Cyber Safeguards
Opus 5 features updated safety guardrails that are approximately 85% less restrictive than those found in Fable 5. The model allows for vulnerability identification within source code but continues to block binary-based scanning, penetration testing, and exploit generation.