Google Gemini 3.8 Flash Performance and Benchmarks
Matthew Bermango watch the original →
the gist
Gemini 3.8 Flash is a highly cost-effective model that performs competitively on software engineering benchmarks like DeepSWE, though it lags behind frontier models in design and complex knowledge work.
Model Performance and Benchmarks
Gemini 3.8 Flash demonstrates strong performance on long-horizon software engineering tasks, achieving a 73.7% score on the DeepSWE V1.1 benchmark, placing it effectively even with Claude Opus 5 and ahead of GPT 5.6 Soul. While it excels in specific agentic and coding tasks, such as the Harvey legal benchmark where it scores 61.4%, it struggles with real-world knowledge work, scoring 1545 on the GDP Val benchmark compared to 1824 for Claude Opus 5. The model also shows promise in agentic computer use, scoring 59% on OSWorld, and maintains high efficiency in cost-per-task metrics.
Specialized Capabilities
Google also introduced Gemini 3.8 Flash Cyber, a variant designed for cybersecurity tasks. This model is restricted to participants in the Fair Wind program and is optimized for vulnerability discovery across 20 programming languages. In internal testing, it significantly outperformed its predecessor, Gemini 3.7 Flash. For general users, the standard 3.8 Flash model shows mixed results in creative and design-heavy tasks. While it successfully generated a functional 3D topographic map of Mount Everest and a basic Doom-style game, it produced simplistic web page designs and lacked the design sophistication of models like Claude Opus 5 for presentation creation.
Pricing and Availability
Gemini 3.8 Flash is positioned as a high-value, low-cost alternative to frontier models. Its introductory pricing is set at $0.75 per million input tokens and $3.75 per million output tokens. These rates are scheduled to increase after the end of the year, though the model remains significantly cheaper than comparable alternatives even at the projected higher price points of $1.50 and $7.50 respectively.