350 output tokens/sec — the fastest model in Google's 3.5 lineup
Huge generational jump over 3.1 Flash-Lite: Terminal-Bench 2.1 54% vs 31%
Punches above its class: SWE-Bench Pro 54.2%, OSWorld-Verified 74.0% at $0.30/$2.50
Weaknesses
Trails full Flash models on hard agentic work (OSWorld 74.0% vs 83.0% for 3.6 Flash)
GPT-5.6 Luna undercuts it on per-token price with stronger benchmark scores
Real-world use cases
What people actually use Gemini 3.5 Flash-Lite for.
Latency-sensitive chat and classification at 350 tokens/sec
Budget agentic pipelines — 54.2% SWE-Bench Pro and computer use built in at $0.30/1M input
Bulk long-context processing with the 1M window at Lite pricing
How Gemini 3.5 Flash-Lite compares
The nearest models people weigh against it, and what actually separates them.
vs Gemini 2.0 Flash — Against Gemini 2.0 Flash (Google), Gemini 3.5 Flash-Lite costs about 82% more per token. Gemini 2.0 Flash is the one to check first if the price difference matters more than the ceiling.
vs Gemini 2.5 Flash — Against Gemini 2.5 Flash (Google), Gemini 3.5 Flash-Lite lands within a few percent on price. Which one wins depends on whether context depth or latency is your constraint.
vs Gemini 3 Flash Preview — Against Gemini 3 Flash Preview (Google), Gemini 3.5 Flash-Lite costs about 38% more per token. Gemini 3 Flash Preview is the one to check first if the price difference matters more than the ceiling.
Price History
Gemini 3.5 Flash-Lite pricing over time
→0% since Aug 7
25 data points · tracked daily since Aug 7, 2026
Ready to try it?
Start using Gemini 3.5 Flash-Lite
High-volume, latency-sensitive workloads at minimal cost. Start free — no card required.
Gemini 2.0 Flash is Google's high-speed, cost-efficient multimodal model built for high-volume production workloads, offering a massive 1M token context window at near-throwaway pricing. It supports text, image, audio, and video inputs with strong instruction-following and tool-use capabilities.
Verdict
The best bang-for-buck multimodal workhorse for developers who need speed, scale, and a massive context window.
Quality score
76%
Pricing
$0.10/1M in
$0.40/1M out
Speed
Very fast
5/5 speed
Context
1.0M tokens
Pricing listed is for standard (non-cached) input/output. Context caching is available and can reduce costs significantly for repeated long-context calls. Image and audio inputs are priced separately. Free tier available via Google AI Studio.
BudgetFastLong ContextMultimodalGoogle
Best for
High-throughput pipelines and agentic tasks where speed and cost matter more than peak reasoning quality.
Gemini 2.5 Flash is Google's fast, cost-efficient multimodal model built for high-throughput tasks requiring a million-token context window at budget pricing. It balances speed and capability across text, code, and vision tasks without the cost of flagship models like Gemini 2.5 Pro.
Verdict
The go-to budget model for long-context and multimodal workloads where speed and scale matter.
Quality score
76%
Pricing
$0.30/1M in
$2.50/1M out
Speed
Very fast
5/5 speed
Context
1.0M tokens
Output cost ($2.5/1M) is disproportionately higher than input cost ($0.3/1M), so generation-heavy use cases may see costs add up faster than expected. Thinking/reasoning mode may be available but incurs additional cost.
BudgetFastLong ContextMultimodalGoogle
Best for
High-volume document processing, summarization, and coding assistance where cost and speed matter more than peak accuracy.
Gemini 3 Flash Preview is Google's budget-tier multimodal model optimized for high-throughput, low-latency tasks at scale. It offers a massive 1M token context window at aggressive pricing, making it a strong contender for cost-sensitive production workloads.
Verdict
A fast, affordable workhorse for long-context and high-volume tasks — just don't build critical systems on a Preview model.
Quality score
74%
Pricing
$0.25/1M in
$1.50/1M out
Speed
Very fast
5/5 speed
Context
1.0M tokens
This is a preview model and may have limited availability, unstable rate limits, and pricing that changes before general availability. Output cost at $3/1M is notably higher than input cost, so applications generating long outputs should budget accordingly.
BudgetLong ContextFastMultimodalPreview
Best for
High-volume document processing, summarization pipelines, and long-context tasks where cost efficiency matters more than frontier-level reasoning.
Gemini 3.5 Flash-Lite costs $0.3 per million input tokens and $2.5 per million output tokens on the API, with cached input at $0.03 per million. A month of 10M input and 2M output tokens runs about $8.00 at list price, before any batch or caching discounts.
What is the context window of Gemini 3.5 Flash-Lite?
Gemini 3.5 Flash-Lite has a 1M tokens context window, with up to 65k tokens of output per response. That is the total of prompt plus response the model can hold in one request.
What is the knowledge cutoff of Gemini 3.5 Flash-Lite?
Gemini 3.5 Flash-Lite's training data runs through March 2026, and the model was released on July 21, 2026. For anything after that date it needs web search or documents in the prompt.
What is Gemini 3.5 Flash-Lite best for?
Gemini 3.5 Flash-Lite is best for high-volume, latency-sensitive workloads at minimal cost. It is a strong fit when that workflow matters more than the tradeoffs around budget pricing and very fast speed.
When should I avoid Gemini 3.5 Flash-Lite?
Pure price-per-benchmark is the criterion — GPT-5.6 Luna wins that math.
What is a cheaper alternative to Gemini 3.5 Flash-Lite?
Gemini 2.0 Flash (Google) at $0.10/1M/1M input against Gemini 3.5 Flash-Lite's $0.30/1M/1M — roughly 82% less per token all in. The best bang-for-buck multimodal workhorse for developers who need speed, scale, and a massive context window. Compare it first if Gemini 3.5 Flash-Lite's pricing is the thing stopping you.
What is a faster alternative to Gemini 3.5 Flash-Lite?
Gemini 2.5 Flash — very fast against Gemini 3.5 Flash-Lite's very fast, with 1.0M tokens of context. Worth the swap when response time is what your users notice rather than the last few points of reasoning depth.
Newsletter
Get notified when Gemini 3.5 Flash-Lite pricing changes
We track pricing daily. When this model drops or spikes, you'll know first.
No spam. Useful updates only. Affiliate disclosures always clearly labeled.