Qwen 3.8 Max Performance and Benchmark Review
AICodeKinggo watch the original →
the gist
Qwen 3.8 Max, a 2.4 trillion parameter model, achieves a 65/80 score on the KingBench benchmark, placing it second overall behind Fable 5 and ahead of Opus 4.8.
Benchmark Performance
Qwen 3.8 Max demonstrates high-level reasoning and coding capabilities, securing a total score of 65 out of 80 on the KingBench suite. The model shows particular strength in agentic workflows, complex mathematical reasoning, and game development tasks, while maintaining competitive performance in frontend and 3D rendering.
Key Task Results
- Bow and Arrow Simulator: Achieved a perfect 10/10 score, successfully implementing physics and leaderboard functionality.
- Mathematical Reasoning: Correctly solved a complex permutation problem with the answer 2460, scoring 10/10.
- Agentic Workflow: Successfully generated a dataset, fine-tuned a Gemma 2B model, and deployed a local web UI autonomously, scoring 10/10.
- 3D Rendering: Scored 8/10 on a Three.js folding table animation and 8/10 on a contact lens case interaction, demonstrating strong spatial and library-specific coding skills.
- Complex Watch Simulation: Scored 3/10 on a 3D wristwatch task, which remains the most difficult benchmark item, though this performance is among the highest recorded for this specific test.
Practical Application
When integrated with Claude Code, the model exhibits high speed and instruction adherence. However, the author notes potential reliability issues during long-running sessions, where the model occasionally fails to create files or complete multi-step harness tasks. The model is currently available as a preview via the Token Plan, Qoder, and QoderWork platforms, with confirmed plans for an open-weight release.