GLM-5.2
Z.ai's MIT-licensed open-weight flagship — the top open-weights coding model of mid-2026, beating GPT-5.5 on agentic coding benchmarks at roughly a sixth of the cost.
Native vision and video, MIT weights, fifteen cents per million.
Cheap multimodal work at scale on MIT-licensed weights
You need top-tier reasoning or a published SWE-bench Verified figure — this is a volume model, not a ceiling model.
Released August 26, 2026. 320B total parameters, 18B active (320B-A18B MoE). Z.ai lists $0.15/1M input, $0.03/1M cached input and $0.50/1M output, with a launch promotion halving those rates through September 9, 2026. Self-reported against GLM-5.2: DeepSWE 63.4 vs 46.2, AutomationBench 48.8 vs 26.2.
$0.15/$0.50 with $0.03 cached input — frontier-adjacent capability at budget-tier pricing
First natively multimodal model in the GLM-5 series: vision and video built in, not bolted on
MIT-licensed weights with a 1M token context and only 18B active parameters per token
No published SWE-bench Verified score
The launch promotion halves these rates only through September 9, 2026 — the $0.15/$0.50 list price applies after that
What people actually use GLM-5.3 Flash for.
High-volume image and video understanding where per-token cost decides the architecture
Self-hosted multimodal pipelines under an MIT license with no commercial restrictions
Bulk coding and automation work — DeepSWE 63.4 against GLM-5.2's 46.2
Cheap multimodal work at scale on MIT-licensed weights. Start free — no card required.
Recommendations are made independently based on real-world use and public benchmarks. See our disclosures for details.
Similar models worth checking before you commit.
Z.ai's MIT-licensed open-weight flagship — the top open-weights coding model of mid-2026, beating GPT-5.5 on agentic coding benchmarks at roughly a sixth of the cost.
Z.ai's newest flagship, aimed squarely at software engineering, autonomous agents and cybersecurity — and the first open-weights model to beat Claude Mythos 5 on a security benchmark.
Google's fastest and most cost-effective 3.5-generation model — low-latency, high-throughput agentic workflows at a fraction of Flash pricing.
GLM-5.3 Flash is best for cheap multimodal work at scale on mit-licensed weights. It is a strong fit when that workflow matters more than the tradeoffs around budget pricing and fast speed.
You need top-tier reasoning or a published SWE-bench Verified figure — this is a volume model, not a ceiling model.
Mistral: Mistral Nemo is the lower-cost option to compare first when you want a similar workflow fit with less token spend.
Gemini 3.5 Flash-Lite is the better pick when response time matters more than maximum depth or premium quality.
Newsletter
We track pricing daily. When this model drops or spikes, you'll know first.
No spam. Useful updates only. Affiliate disclosures always clearly labeled.
No reviews yet — be the first.