Gemini 3.5 Flash-Lite
Google's fastest and most cost-effective 3.5-generation model — low-latency, high-throughput agentic workflows at a fraction of Flash pricing.
Native vision and video, MIT weights, fifteen cents per million.
Cheap multimodal work at scale on MIT-licensed weights
You need top-tier reasoning or a published SWE-bench Verified figure — this is a volume model, not a ceiling model.
Compare every model's knowledge cutoff, max output, and context window.
Released August 26, 2026. 320B total parameters, 18B active (320B-A18B MoE). Z.ai lists $0.15/1M input, $0.03/1M cached input and $0.50/1M output, with a launch promotion halving those rates through September 9, 2026. Self-reported against GLM-5.2: DeepSWE 63.4 vs 46.2, AutomationBench 48.8 vs 26.2.
$0.15/$0.50 with $0.03 cached input — frontier-adjacent capability at budget-tier pricing
First natively multimodal model in the GLM-5 series: vision and video built in, not bolted on
MIT-licensed weights with a 1M token context and only 18B active parameters per token
No published SWE-bench Verified score
The launch promotion halves these rates only through September 9, 2026 — the $0.15/$0.50 list price applies after that
What people actually use GLM-5.3 Flash for.
High-volume image and video understanding where per-token cost decides the architecture
Self-hosted multimodal pipelines under an MIT license with no commercial restrictions
Bulk coding and automation work — DeepSWE 63.4 against GLM-5.2's 46.2
The nearest models people weigh against it, and what actually separates them.
vs Gemini 3.5 Flash-Lite — Against Gemini 3.5 Flash-Lite (Google), GLM-5.3 Flash runs about 77% cheaper per token, gives up 1x on context and answers slower. Which one wins depends on whether context depth or latency is your constraint.
vs GLM-5.2 — Against GLM-5.2 (Z.ai), GLM-5.3 Flash runs about 89% cheaper per token and answers faster. Take GLM-5.3 Flash unless you specifically need what GLM-5.2 does better.
vs GLM-5.3 — Against GLM-5.3 (Z.ai), GLM-5.3 Flash runs about 89% cheaper per token and answers faster. Take GLM-5.3 Flash unless you specifically need what GLM-5.3 does better.
Price History
→0% since Sep 1
7 data points · tracked daily since Sep 1, 2026
Cheap multimodal work at scale on MIT-licensed weights. Start free — no card required.
Recommendations are made independently based on real-world use and public benchmarks. See our disclosures for details.
Similar models worth checking before you commit.
Google's fastest and most cost-effective 3.5-generation model — low-latency, high-throughput agentic workflows at a fraction of Flash pricing.
Z.ai's MIT-licensed open-weight flagship — the top open-weights coding model of mid-2026, beating GPT-5.5 on agentic coding benchmarks at roughly a sixth of the cost.
Z.ai's newest flagship, aimed squarely at software engineering, autonomous agents and cybersecurity — and the first open-weights model to beat Claude Mythos 5 on a security benchmark.
GLM-5.3 Flash costs $0.15 per million input tokens and $0.5 per million output tokens on the API. A month of 10M input and 2M output tokens runs about $2.50 at list price, before any batch or caching discounts.
GLM-5.3 Flash has a 1M tokens context window, with up to 131k tokens of output per response. That is the total of prompt plus response the model can hold in one request.
GLM-5.3 Flash is best for cheap multimodal work at scale on mit-licensed weights. It is a strong fit when that workflow matters more than the tradeoffs around budget pricing and fast speed.
You need top-tier reasoning or a published SWE-bench Verified figure — this is a volume model, not a ceiling model.
Mistral Small 3.1 (Mistral) at $0.10/1M/1M input against GLM-5.3 Flash's $0.15/1M/1M — roughly 38% less per token all in. Ultra-cheap multimodal model for massive-volume, low-complexity pipelines. Compare it first if GLM-5.3 Flash's pricing is the thing stopping you.
Gemini 3.5 Flash-Lite — very fast against GLM-5.3 Flash's fast, with 1.0M tokens of context. Worth the swap when response time is what your users notice rather than the last few points of reasoning depth.
Newsletter
We track pricing daily. When this model drops or spikes, you'll know first.
No spam. Useful updates only. Affiliate disclosures always clearly labeled.
No reviews yet — be the first.