Gemini 3.5 Flash
Google's I/O 2026 headliner — a Flash-tier model that beats Gemini 3.1 Pro on agentic and coding benchmarks while running roughly 4x faster than comparable frontier models.
Best for research and deep document analysis — 2M context at the best premium price.
Research, deep document analysis, and long-context reasoning at competitive pricing
Your primary use case is writing quality or agentic coding — Claude wins both.
Compare every model's knowledge cutoff, max output, and context window.
The 2M context window is a genuine competitive advantage — no other frontier model gets close for document-heavy workflows.
2M token context window — the largest of any frontier model
Leads ARC-AGI-2 reasoning benchmark at 77.1%
Best price-to-performance among premium models at $2/$12 per 1M tokens
Slower than Flash for everyday lightweight tasks
Claude Sonnet 4.6 is better for writing quality
What people actually use Gemini 3.1 Pro for.
Analyzing entire contracts, codebases, or research corpora in a single 2M-token prompt
Due diligence synthesis across large sets of financial documents or legal agreements
Multi-step reasoning across dense technical specifications with Deep Think mode
The nearest models people weigh against it, and what actually separates them.
vs Gemini 3.5 Flash — Against Gemini 3.5 Flash (Google), Gemini 3.1 Pro costs about 25% more per token, takes 1.9x the context and answers slower. Gemini 3.5 Flash is the one to check first if the price difference matters more than the ceiling.
vs Gemini 3.6 Flash — Against Gemini 3.6 Flash (Google), Gemini 3.1 Pro costs about 68% more per token, takes 1.9x the context and answers slower. Gemini 3.6 Flash is the one to check first if the price difference matters more than the ceiling.
vs GPT-6 Astra — Against GPT-6 Astra (OpenAI), Gemini 3.1 Pro runs about 77% cheaper per token, takes 1.9x the context and answers faster. Take Gemini 3.1 Pro unless you specifically need what GPT-6 Astra does better.
Price History
→0% since May 17
90 data points · tracked daily since May 17, 2026
Research, deep document analysis, and long-context reasoning at competitive pricing. Start free — no card required.
Recommendations are made independently based on real-world use and public benchmarks. See our disclosures for details.
Similar models worth checking before you commit.
Google's I/O 2026 headliner — a Flash-tier model that beats Gemini 3.1 Pro on agentic and coding benchmarks while running roughly 4x faster than comparable frontier models.
Google's efficiency-focused successor to 3.5 Flash — higher scores on every benchmark Google tested, ~17% fewer output tokens, and cheaper output pricing.
OpenAI's September 3, 2026 frontier release — the first GPT-6 model and OpenAI's answer to Claude Fable 5.1 two days earlier. State of the art on computer use (OSWorld 2.0 72.6% in ~47% less time than GPT-5.6 Sol), agentic coding (Terminal-Bench 4.0 57.9%), and frontier math (FrontierMath Tier 4 97.6%). $10/$50 per 1M tokens, 1.05M context, 128K output, knowledge cutoff April 30, 2026.
Gemini 3.1 Pro costs $2 per million input tokens and $12 per million output tokens on the API, with cached input at $0.2 per million. A month of 10M input and 2M output tokens runs about $44.00 at list price, before any batch or caching discounts.
Gemini 3.1 Pro has a 1M tokens context window, with up to 64k tokens of output per response. That is the total of prompt plus response the model can hold in one request.
Gemini 3.1 Pro's training data runs through January 2025, and the model was released on February 19, 2026. For anything after that date it needs web search or documents in the prompt.
Gemini 3.1 Pro is best for research, deep document analysis, and long-context reasoning at competitive pricing. It is a strong fit when that workflow matters more than the tradeoffs around premium pricing and balanced speed.
Your primary use case is writing quality or agentic coding — Claude wins both.
Gemini 3.6 Flash (Google) at $0.75/1M/1M input against Gemini 3.1 Pro's $2.00/1M/1M — roughly 68% less per token all in. Best Gemini for agents — efficiency king with native computer use. Compare it first if Gemini 3.1 Pro's pricing is the thing stopping you.
Gemini 3.5 Flash — fast against Gemini 3.1 Pro's balanced, with 1.0M tokens of context. Worth the swap when response time is what your users notice rather than the last few points of reasoning depth.
Newsletter
We track pricing daily. When this model drops or spikes, you'll know first.
No spam. Useful updates only. Affiliate disclosures always clearly labeled.
No reviews yet — be the first.