DeepSeek V4-Flash
A 284B-parameter (13B active) MoE workhorse re-post-trained for agentic and coding tasks — beats the V4-Pro preview on every published agent benchmark at ultra-commodity pricing.
SWE-bench Pro 62.5 at sixteen cents per million input.
Cheap high-throughput coding and reasoning
You need Alibaba's maximum capability — that is Qwen 3.8 Max — or a SWE-bench Verified number.
Released August 26, 2026. The open-weight release is Qwen3.8-Flash-Next, a preview of the Qwen4 architecture: 125B mixture-of-experts with 6B active per token, a 51B n-gram embedding table and a 4B multi-token prediction layer. Qwen 3.8 Flash is the production API version on Qwen Cloud at $0.16/$0.47.
SWE-bench Pro 62.5 — competitive with models several times its price
Only 6B active parameters per token from a 125B mixture-of-experts, so throughput is high and hosting is cheap
991K context window at $0.16/$0.47
No published SWE-bench Verified score, only SWE-bench Pro
An architecture preview rather than a settled flagship — Qwen 3.8 Max remains Alibaba's top-end model
What people actually use Qwen 3.8 Flash for.
Volume coding work where SWE-bench Pro 62.5 is enough and cost per token dominates
Near-1M-context document processing at budget-tier rates
Self-hosted inference on modest hardware thanks to 6B active parameters per token
Cheap high-throughput coding and reasoning. Start free — no card required.
Recommendations are made independently based on real-world use and public benchmarks. See our disclosures for details.
Similar models worth checking before you commit.
A 284B-parameter (13B active) MoE workhorse re-post-trained for agentic and coding tasks — beats the V4-Pro preview on every published agent benchmark at ultra-commodity pricing.
DeepSeek's 1.6T-parameter (49B active) MoE flagship with hybrid sparse attention — near-frontier coding and reasoning at roughly a tenth of closed-rival pricing, MIT-licensed open weights.
Z.ai's MIT-licensed open-weight flagship — the top open-weights coding model of mid-2026, beating GPT-5.5 on agentic coding benchmarks at roughly a sixth of the cost.
Qwen 3.8 Flash is best for cheap high-throughput coding and reasoning. It is a strong fit when that workflow matters more than the tradeoffs around budget pricing and very fast speed.
You need Alibaba's maximum capability — that is Qwen 3.8 Max — or a SWE-bench Verified number.
Mistral: Mistral Nemo is the lower-cost option to compare first when you want a similar workflow fit with less token spend.
DeepSeek V4-Flash is the better pick when response time matters more than maximum depth or premium quality.
Newsletter
We track pricing daily. When this model drops or spikes, you'll know first.
No spam. Useful updates only. Affiliate disclosures always clearly labeled.
No reviews yet — be the first.