Theo's AI Model Tier List: Performance, Cost, and Utility
Theo - t3.gggo watch the original →
the gist
A practical, developer-focused ranking of current AI models, emphasizing that 'best' is highly context-dependent, with specific focus on token efficiency, vision capabilities, and real-world coding utility.
The Fallacy of the 'Best' Model
Theo argues that a single tier list is inherently flawed because model performance varies wildly across different axes: coding capability, token efficiency, speed, and specific task-based utility (e.g., summarization vs. complex architectural rewrites). The "best" model is often determined by the specific constraints of the workflow—whether you are optimizing for cost, latency, or the ability to handle complex, multi-step reasoning.
The Hierarchy of Utility
The ranking prioritizes models that provide tangible value in day-to-day development.
- S-Tier (The Heavy Hitters): Models like 56 Soul stand out for their ability to handle long-running, complex coding tasks (like full app rewrites in SwiftUI) with high intelligence.
- A-Tier (The Workhorses): Models like Luna are praised not necessarily for raw intelligence, but for their speed, low cost, and reliability in agentic workflows like title generation, data categorization, and managing thread context.
- B/C-Tier (The Niche Players): Models like Kimmy K3 are recognized for impressive open-weight performance and design taste, but are held back by pricing structures and licensing constraints that make them less economical than proprietary alternatives for high-volume use.
- D/F-Tier (The Disappointments): Models that lack essential features like vision (e.g., DeepSeek V4 Pro, GLM53) or those that suffer from regression in efficiency (e.g., Grok 46) are penalized heavily, regardless of their underlying architecture.
The Hidden Costs of 'Open Weight'
A recurring theme is the misconception that open-weight models are always cheaper. Theo highlights that licensing agreements (like those for Kimmy K3) often mandate revenue-sharing or pricing floors, meaning they don't always undercut proprietary models on price. Furthermore, token efficiency is a critical, often overlooked metric; a model that is cheaper per token but less efficient at utilizing those tokens can end up being more expensive for a total task than a "smarter" model that requires fewer tokens to reach the same conclusion.
The Vision Requirement
Vision capabilities have become a non-negotiable requirement for modern developer workflows. Models that lack vision are relegated to lower tiers because the ability to paste images for context is essential for debugging and UI/UX iteration. The lack of vision in high-end, "Pro" tier models is described as "embarrassing" in the current landscape.