Grok 4.6 Performance and xAI's Coding Strategy
Matthew Bermango watch the original →
the gist
xAI's Grok 4.6 shows significant gains in coding and legal benchmarks, driven by recursive training on Grok 4.5 outputs and the integration of Cursor's proprietary coding data.
Performance and Benchmarks
Grok 4.6 represents an iterative improvement over the 4.5 model, specifically targeting coding and knowledge work. It achieves top-tier results in specialized benchmarks, notably leading the Harvey Lab legal benchmark with a 15.8% score compared to 11.3% for Fable 5. In coding-specific evaluations, it reached a 26% score on Terminal Bench, up from 15% in the previous version. While it remains competitive in intelligence indices, it occupies a higher cost-per-task bracket than its predecessor, moving from approximately 36 cents to 83 cents per task. Despite these gains, qualitative testing in UI generation shows it still trails GPT 5.6 in visual layout and spacing accuracy.
The Data-Compute Flywheel
The current trajectory of xAI is defined by the synthesis of Cursor's coding data and xAI's massive GPU infrastructure. The company utilized Grok 4.5 to curate and regenerate SFT (Supervised Fine-Tuning) trajectories across reasoning and software engineering domains to train Grok 4.6. This strategy mirrors the successful development loops seen at Anthropic and OpenAI, where high-quality coding data is used to iteratively improve the model's own reasoning capabilities. The integration of this model into Grokbot and Cursor suggests a dual-pronged approach: maintaining a developer-focused IDE experience while expanding into a simplified, artifact-based interface for non-technical users.