Gemini 3.5 Flash-Lite
Google's fastest and most cost-effective 3.5-generation model — low-latency, high-throughput agentic workflows at a fraction of Flash pricing.
A fast, affordable workhorse for long-context and high-volume tasks — just don't build critical systems on a Preview model.
High-volume document processing, summarization pipelines, and long-context tasks where cost efficiency matters more than frontier-level reasoning.
You need reliable, stable API access for production applications or require strong multi-step reasoning and complex instruction adherence.
This is a preview model and may have limited availability, unstable rate limits, and pricing that changes before general availability. Output cost at $3/1M is notably higher than input cost, so applications generating long outputs should budget accordingly.
Massive 1M token context window enables ingestion of entire codebases or lengthy legal documents in a single call
Very competitive $0.5/$3 per million token pricing undercuts GPT-4.1 mini and Claude Haiku 3.5 for long-context use
Fast inference speed suitable for real-time applications and batch processing pipelines
Multimodal input support inherited from Gemini architecture handles images, text, and code natively
Preview status means API stability, rate limits, and pricing are subject to change without notice
Complex multi-step reasoning and nuanced instruction-following lag behind Gemini 3 Pro and Claude Sonnet 4.6
Output cost at $3/1M tokens is higher relative to input, making verbose-output tasks less economical
What people actually use Google: Gemini 3 Flash Preview for.
Summarizing entire 500-page PDF reports or legal contracts in a single API call using the 1M token context
Building a cost-efficient chatbot or Q&A pipeline over a large knowledge base at scale
Batch-processing thousands of multimodal inputs (images + text) for classification or tagging workflows
Price History
↓50% since May 9
88 data points · tracked daily since May 9, 2026
High-volume document processing, summarization pipelines, and long-context tasks where cost efficiency matters more than frontier-level reasoning.. Start free — no card required.
Recommendations are made independently based on real-world use and public benchmarks. See our disclosures for details.
Similar models worth checking before you commit.
Google's fastest and most cost-effective 3.5-generation model — low-latency, high-throughput agentic workflows at a fraction of Flash pricing.
Gemma 4 26B A4B is a sparse mixture-of-experts open model from Google, activating only ~4B parameters per forward pass despite having 26B total parameters. It offers a 262K context window at budget pricing, making it one of the more capable open-weight models for its cost tier.
Gemma 4 31B is Google's open-weight instruction-tuned model offering a strong balance of capability and cost efficiency at just $0.14/$0.40 per million tokens. It features a 262K context window and is designed for developers who need capable on-premise or API-hosted inference without flagship pricing.
Pricing moves, ranking shifts, and capability updates.
Google: Gemini 3 Flash Preview output pricing changed from $3.00/1M to $1.50/1M (↓ cheaper, 50% cut).
View modelGoogle: Gemini 3 Flash Preview input pricing changed from $0.50/1M to $0.25/1M (↓ cheaper, 50% cut).
View modelGoogle: Gemini 3 Flash Preview output pricing changed from $1.50/1M to $3.00/1M (↑ more expensive, 100% increase).
View modelGoogle: Gemini 3 Flash Preview input pricing changed from $0.25/1M to $0.50/1M (↑ more expensive, 100% increase).
View modelGoogle: Gemini 3 Flash Preview output pricing changed from $3.00/1M to $1.50/1M (↓ cheaper, 50% cut).
View modelGoogle: Gemini 3 Flash Preview input pricing changed from $0.50/1M to $0.25/1M (↓ cheaper, 50% cut).
View modelGoogle: Gemini 3 Flash Preview (Google) is now indexed. A fast, affordable workhorse for long-context and high-volume tasks — just don't build critical systems on a Preview model.
View modelGoogle: Gemini 3 Flash Preview is best for high-volume document processing, summarization pipelines, and long-context tasks where cost efficiency matters more than frontier-level reasoning.. It is a strong fit when that workflow matters more than the tradeoffs around budget pricing and very fast speed.
You need reliable, stable API access for production applications or require strong multi-step reasoning and complex instruction adherence.
Mistral: Mistral Nemo is the lower-cost option to compare first when you want a similar workflow fit with less token spend.
Gemini 3.5 Flash-Lite is the better pick when response time matters more than maximum depth or premium quality.
Newsletter
We track pricing daily. When this model drops or spikes, you'll know first.
No spam. Useful updates only. Affiliate disclosures always clearly labeled.
No reviews yet — be the first.