Gemini 3.5 Flash
Google's I/O 2026 headliner — a Flash-tier model that beats Gemini 3.1 Pro on agentic and coding benchmarks while running roughly 4x faster than comparable frontier models.
Google's fast agentic workhorse — strong coding at Flash pricing.
Fast, low-cost agentic coding and multimodal work, including video input
You need long single responses (65K output cap) or your workload is output-heavy enough that its verbosity erases the per-token saving.
Compare every model's knowledge cutoff, max output, and context window.
Released September 2, 2026, Google's third Flash release in about six weeks (3.6 Flash July 21, 3.7 Flash August 13). Gateway id google/gemini-3.8-flash. $0.75/$3.75 per 1M; Flex $0.375/$1.875. 1,000,000 context, 65,535 max output; text, image, PDF and video input. Google launch figures: Terminal-Bench 2.1 90.8 (3.7 Flash 81.6); HLE-Verified 54.9. A Gemini 3.8 Flash Cyber variant is restricted to governments and vetted partners. Verified October 10, 2026.
Same $0.75/$3.75 price as 3.7 Flash with Google-reported gains on coding and agent benchmarks
Artificial Analysis measured about 302 output tokens per second, among the fastest models it tracks
Accepts text, images, PDF and video natively
Verbose: Artificial Analysis needed 120M output tokens to run its index against a 71M median, so cost per task runs above the sticker price
65K max output — half of what Claude and OpenAI's current models allow
What people actually use Gemini 3.8 Flash for.
Terminal and coding agents — Google reports 90.8% on Terminal-Bench 2.1, up from 81.6% for 3.7 Flash
Video, image and PDF understanding in one 1M-context call
Finance and legal agent workflows, where Google reports gains on Vals Finance Agent V2 and Harvey's legal benchmark
The nearest models people weigh against it, and what actually separates them.
vs Gemini 3.5 Flash — Against Gemini 3.5 Flash (Google), Gemini 3.8 Flash runs about 57% cheaper per token, gives up 1x on context and answers faster. Take Gemini 3.8 Flash unless you specifically need what Gemini 3.5 Flash does better.
vs Gemini 3.6 Flash — Against Gemini 3.6 Flash (Google), Gemini 3.8 Flash lands within a few percent on price, gives up 1x on context and answers faster. Which one wins depends on whether context depth or latency is your constraint.
vs Muse Spark 1.3 — Against Muse Spark 1.3 (Meta), Gemini 3.8 Flash runs about 18% cheaper per token, gives up 1x on context and answers faster. Take Gemini 3.8 Flash unless you specifically need what Muse Spark 1.3 does better.
Fast, low-cost agentic coding and multimodal work, including video input. Start free — no card required.
Recommendations are made independently based on real-world use and public benchmarks. See our disclosures for details.
Similar models worth checking before you commit.
Google's I/O 2026 headliner — a Flash-tier model that beats Gemini 3.1 Pro on agentic and coding benchmarks while running roughly 4x faster than comparable frontier models.
Google's efficiency-focused successor to 3.5 Flash — higher scores on every benchmark Google tested, ~17% fewer output tokens, and cheaper output pricing.
Meta's September 2, 2026 update to Muse Spark — a multimodal reasoning model for long-horizon agent and coding work, with a 1M context and native understanding of video, images and documents.
Pricing moves, ranking shifts, and capability updates.
Google released Gemini 3.8 Flash on September 2, 2026, its third Flash model in about six weeks, at the same $0.75/$3.75 per million tokens as 3.7 Flash. Google reports 90.8% on Terminal-Bench 2.1 (3.7 Flash: 81.6%) and 54.9% on HLE-Verified. Artificial Analysis measured about 302 output tokens per second but also found it verbose, using 120M output tokens to run its index against a 71M median — so cost per task runs above the per-token price. It accepts text, images, PDF and video, with a 1M-token context and 65K output. A Gemini 3.8 Flash Cyber variant is limited to governments and vetted partners. Verified October 10, 2026.
View modelGemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens on the API, with cached input at $0.075 per million. A month of 10M input and 2M output tokens runs about $15.00 at list price, before any batch or caching discounts.
Gemini 3.8 Flash has a 1M tokens context window, with up to 66k tokens of output per response. That is the total of prompt plus response the model can hold in one request.
Gemini 3.8 Flash is best for fast, low-cost agentic coding and multimodal work, including video input. It is a strong fit when that workflow matters more than the tradeoffs around balanced pricing and very fast speed.
You need long single responses (65K output cap) or your workload is output-heavy enough that its verbosity erases the per-token saving.
Claude Sonnet 5.5 (Anthropic) at $2.00/1M/1M input against Gemini 3.8 Flash's $0.75/1M/1M. Near-Opus 5.5 quality on scoped work at half the price. Compare it first if Gemini 3.8 Flash's pricing is the thing stopping you.
Gemini 3.5 Flash — fast against Gemini 3.8 Flash's very fast, with 1.0M tokens of context. Worth the swap when response time is what your users notice rather than the last few points of reasoning depth.
Newsletter
We track pricing daily. When this model drops or spikes, you'll know first.
No spam. Useful updates only. Affiliate disclosures always clearly labeled.
No reviews yet — be the first.