UseRightAI
HomeModelsAsk AIComparePricingWhat's New
UseRightAICut through AI hype. Pick what works.

Independent AI model tracker. Live pricing, real benchmarks, zero vendor bias.

X (Twitter)LinkedInUpdatesContact

Compare

Opus 4.8 vs Opus 4.7Fable 5 vs Opus 4.8New AI Models 2026ChatGPT vs ClaudeGPT-4o vs Claude SonnetClaude vs GeminiDeepSeek vs ChatGPTMistral vs ClaudeGemini Flash vs GPT-4o MiniLlama vs ChatGPTAll comparisons →Build your own →

Best For

CodingWritingDevelopersProduct ManagersDesignersSalesBest Cheap AIBest Free AI

Pricing & Data

API Token PricingCost per TaskPrice HistoryBenchmark ScoresPrivacy & SafetySubscription PlansPlan Usage LimitsCost CalculatorWhich AI is Cheapest?Cheapest AI APIs

Company

About UseRightAIContactWhat ChangedAll ModelsEditorial PolicyDisclosuresPrivacy PolicyTerms of Service

© 2026 UseRightAI. Independent · Free forever · Not affiliated with any AI provider.

Affiliate links are clearly labeled. See disclosures.

HomeModelsGemini 3.5 Flash-Lite
GoogleBudget

Gemini 3.5 Flash-Lite

Fastest budget multimodal model — 350 tokens/sec at Lite pricing.

78
Coding
76
Writing
78
Research
83
Images
96
Value
84
Long Context
Use this when

High-volume, latency-sensitive workloads at minimal cost

Skip this if

Pure price-per-benchmark is the criterion — GPT-5.6 Luna wins that math.

Pricing
$0.30/1M in
$2.50/1M out
→0%since Aug 2026
Context
1.0M tokens
Speed
Very fast

Gemini 3.5 Flash-Litespecs & pricing

Verified Sep 4, 2026 against the AI Gateway catalog
Input price
$0.30 / 1M tokens
Output price
$2.50 / 1M tokens
Cached input(prompt-cache read)
$0.030 / 1M tokens
Context window
1M tokens
Max output
65k tokens
Knowledge cutoff
Mar 2026
Released
Jul 21, 2026
Input modalities
Text, Image, PDF, Video
Output modalities
Text
Reasoning mode
Yes
Tool use
Yes
Gateway model ID
google/gemini-3.5-flash-lite

Compare every model's knowledge cutoff, max output, and context window.

Released July 21, 2026. Batch pricing $0.15/$1.25; cached input $0.03/1M. Rolling out to Google Search.

How to access
Subscription
Google One AI Premium — $19.99/mo
API
$0.3/1M input tokens
Subscription = chat interface. API = build with it. Compare all subscription plans · Google One AI Premium usage limits
Switch to instead if...
Best overall
GPT-6 Astra
Cheaper option
Gemini 2.0 Flash
Faster option
Gemini 2.5 Flash

Strengths

350 output tokens/sec — the fastest model in Google's 3.5 lineup

Huge generational jump over 3.1 Flash-Lite: Terminal-Bench 2.1 54% vs 31%

Punches above its class: SWE-Bench Pro 54.2%, OSWorld-Verified 74.0% at $0.30/$2.50

Weaknesses

Trails full Flash models on hard agentic work (OSWorld 74.0% vs 83.0% for 3.6 Flash)

GPT-5.6 Luna undercuts it on per-token price with stronger benchmark scores

Real-world use cases

What people actually use Gemini 3.5 Flash-Lite for.

Latency-sensitive chat and classification at 350 tokens/sec

Budget agentic pipelines — 54.2% SWE-Bench Pro and computer use built in at $0.30/1M input

Bulk long-context processing with the 1M window at Lite pricing

How Gemini 3.5 Flash-Lite compares

The nearest models people weigh against it, and what actually separates them.

vs Gemini 2.0 Flash — Against Gemini 2.0 Flash (Google), Gemini 3.5 Flash-Lite costs about 82% more per token. Gemini 2.0 Flash is the one to check first if the price difference matters more than the ceiling.

vs Gemini 2.5 Flash — Against Gemini 2.5 Flash (Google), Gemini 3.5 Flash-Lite lands within a few percent on price. Which one wins depends on whether context depth or latency is your constraint.

vs Gemini 3 Flash Preview — Against Gemini 3 Flash Preview (Google), Gemini 3.5 Flash-Lite costs about 38% more per token. Gemini 3 Flash Preview is the one to check first if the price difference matters more than the ceiling.

Price History

Gemini 3.5 Flash-Lite pricing over time

→0% since Aug 7

$0.324$0.312$0.300$0.288$0.276Aug 7Aug 12Aug 17Aug 21Sep 2Sep 7

25 data points · tracked daily since Aug 7, 2026

Ready to try it?

Start using Gemini 3.5 Flash-Lite

High-volume, latency-sensitive workloads at minimal cost. Start free — no card required.

Try Gemini 3.5 Flash-Lite freeCompare alternatives

Recommendations are made independently based on real-world use and public benchmarks. See our disclosures for details.

Compare alternatives

Similar models worth checking before you commit.

All Gemini 3.5 Flash-Lite alternatives →
GoogleBudget

Gemini 2.0 Flash

Gemini 2.0 Flash is Google's high-speed, cost-efficient multimodal model built for high-volume production workloads, offering a massive 1M token context window at near-throwaway pricing. It supports text, image, audio, and video inputs with strong instruction-following and tool-use capabilities.

Verdict
The best bang-for-buck multimodal workhorse for developers who need speed, scale, and a massive context window.
Quality score
76%
Pricing
$0.10/1M in
$0.40/1M out
Speed
Very fast
5/5 speed
Context
1.0M tokens
Pricing listed is for standard (non-cached) input/output. Context caching is available and can reduce costs significantly for repeated long-context calls. Image and audio inputs are priced separately. Free tier available via Google AI Studio.
BudgetFastLong ContextMultimodalGoogle
Best for
High-throughput pipelines and agentic tasks where speed and cost matter more than peak reasoning quality.
View model
GoogleBudget

Gemini 2.5 Flash

Gemini 2.5 Flash is Google's fast, cost-efficient multimodal model built for high-throughput tasks requiring a million-token context window at budget pricing. It balances speed and capability across text, code, and vision tasks without the cost of flagship models like Gemini 2.5 Pro.

Verdict
The go-to budget model for long-context and multimodal workloads where speed and scale matter.
Quality score
76%
Pricing
$0.30/1M in
$2.50/1M out
Speed
Very fast
5/5 speed
Context
1.0M tokens
Output cost ($2.5/1M) is disproportionately higher than input cost ($0.3/1M), so generation-heavy use cases may see costs add up faster than expected. Thinking/reasoning mode may be available but incurs additional cost.
BudgetFastLong ContextMultimodalGoogle
Best for
High-volume document processing, summarization, and coding assistance where cost and speed matter more than peak accuracy.
View model
GoogleBudget

Gemini 3 Flash Preview

Gemini 3 Flash Preview is Google's budget-tier multimodal model optimized for high-throughput, low-latency tasks at scale. It offers a massive 1M token context window at aggressive pricing, making it a strong contender for cost-sensitive production workloads.

Verdict
A fast, affordable workhorse for long-context and high-volume tasks — just don't build critical systems on a Preview model.
Quality score
74%
Pricing
$0.25/1M in
$1.50/1M out
Speed
Very fast
5/5 speed
Context
1.0M tokens
This is a preview model and may have limited availability, unstable rate limits, and pricing that changes before general availability. Output cost at $3/1M is notably higher than input cost, so applications generating long outputs should budget accordingly.
BudgetLong ContextFastMultimodalPreview
Best for
High-volume document processing, summarization pipelines, and long-context tasks where cost efficiency matters more than frontier-level reasoning.
View model

Gemini 3.5 Flash-Lite head-to-head

All Gemini 3.5 Flash-Lite alternatives →GPT-5.6 Luna vs Gemini 3.5 Flash-Lite →GLM-5.3 Flash vs Gemini 3.5 Flash-Lite →View benchmark scores →

FAQ

How much does Gemini 3.5 Flash-Lite cost?

Gemini 3.5 Flash-Lite costs $0.3 per million input tokens and $2.5 per million output tokens on the API, with cached input at $0.03 per million. A month of 10M input and 2M output tokens runs about $8.00 at list price, before any batch or caching discounts.

What is the context window of Gemini 3.5 Flash-Lite?

Gemini 3.5 Flash-Lite has a 1M tokens context window, with up to 65k tokens of output per response. That is the total of prompt plus response the model can hold in one request.

What is the knowledge cutoff of Gemini 3.5 Flash-Lite?

Gemini 3.5 Flash-Lite's training data runs through March 2026, and the model was released on July 21, 2026. For anything after that date it needs web search or documents in the prompt.

What is Gemini 3.5 Flash-Lite best for?

Gemini 3.5 Flash-Lite is best for high-volume, latency-sensitive workloads at minimal cost. It is a strong fit when that workflow matters more than the tradeoffs around budget pricing and very fast speed.

When should I avoid Gemini 3.5 Flash-Lite?

Pure price-per-benchmark is the criterion — GPT-5.6 Luna wins that math.

What is a cheaper alternative to Gemini 3.5 Flash-Lite?

Gemini 2.0 Flash (Google) at $0.10/1M/1M input against Gemini 3.5 Flash-Lite's $0.30/1M/1M — roughly 82% less per token all in. The best bang-for-buck multimodal workhorse for developers who need speed, scale, and a massive context window. Compare it first if Gemini 3.5 Flash-Lite's pricing is the thing stopping you.

What is a faster alternative to Gemini 3.5 Flash-Lite?

Gemini 2.5 Flash — very fast against Gemini 3.5 Flash-Lite's very fast, with 1.0M tokens of context. Worth the swap when response time is what your users notice rather than the last few points of reasoning depth.

Newsletter

Get notified when Gemini 3.5 Flash-Lite pricing changes

We track pricing daily. When this model drops or spikes, you'll know first.

No spam. Useful updates only. Affiliate disclosures always clearly labeled.

User reviews

No reviews yet — be the first.