Claude Opus 5 Performance Review
AICodeKinggo watch the original →
the gist
Claude Opus 5 excels at agentic coding and complex reasoning but shows regressions in 3D rendering and general assistant capabilities compared to its predecessor.
Reasoning and Agentic Performance
Claude Opus 5 demonstrates high proficiency in logic-heavy and long-horizon tasks. It achieved perfect scores on the elevator simulation, bow and arrow game, math permutation problems, and a complex autonomous finetuning task involving MLX and local web UI generation. The model shows strong alignment and reasoning capabilities, consistently hitting 10 out of 10 on tasks requiring multi-step planning and code execution.
Visual and 3D Limitations
Despite its reasoning gains, the model underperforms in visual and 3D spatial tasks. It regressed on the folding table animation task, scoring a 5 out of 10 compared to the 8 out of 10 achieved by Opus 4.8. Other visual benchmarks, such as the 3D wristwatch and SVG generation, remain stagnant, indicating that the model's training focus prioritized agentic workflows over spatial or visual fidelity.
Real-World Usability
In practical application, the model exhibits high verbosity and unnecessary file-system exploration, often touching more files than required for a given task. This behavior increases token consumption and latency. General knowledge and conversational capabilities feel less capable than Fable 5 or GPT-5.6 Sol, suggesting that while Opus 5 is a cost-effective agentic workhorse, it is not a direct upgrade for general-purpose assistant tasks.