Anthropic Claude Opus 5 Performance and Cost Analysis

Matthew Bermango watch the original →

Claude Opus 5 outperforms Fable 5 on most benchmarks while costing half as much, marking a significant shift toward prioritizing cost-per-task efficiency over raw token pricing.

Performance and Efficiency Gains

Claude Opus 5 demonstrates superior performance across major benchmarks compared to its predecessor, Opus 4.8, and the Fable 5 model. Notably, the model achieved a 30% score on the ARC-AGI 3 benchmark, a significant jump from previous state-of-the-art results which hovered around 8%. While Opus 5 shows improved capabilities in coding and automation, it exhibits a deliberate reduction in cyber-security exploit development compared to the unrestricted Mythos 5 model, suggesting Anthropic implemented strict safety guardrails without degrading general reasoning performance.

The Shift to Cost-Per-Task Metrics

The primary value proposition of Opus 5 is its efficiency, priced at $5 per million input tokens and $25 per million output tokens. The author emphasizes that evaluating models based solely on price-per-token is misleading, as some models require significantly more tokens to complete the same task. Opus 5 consistently ranks as the most efficient option when measured by cost-per-task, outperforming GPT 5.6 Soul in both accuracy and total expenditure for complex enterprise workflows like due diligence and data analysis.

Operational Constraints

Despite the performance improvements, users may encounter automatic API fallbacks to Opus 4.8 if the safety classifiers flag a request. This behavior mirrors existing patterns in other frontier models, where aggressive safety filtering can interrupt workflows. The author notes that while Opus 5 is a strong daily driver, it may still be subject to government-mandated usage restrictions due to its high reasoning capabilities.

  • #ai
  • #llm
  • #benchmarking

summary by google/gemini-3.1-flash-lite. probably wrong about something. check the source.