DeepSeek V4-Flash
A 284B-parameter (13B active) MoE workhorse re-post-trained for agentic and coding tasks — beats the V4-Pro preview on every published agent benchmark at ultra-commodity pricing.
SWE-bench Pro 62.5 at sixteen cents per million input.
Cheap high-throughput coding and reasoning
You need Alibaba's maximum capability — that is Qwen 3.8 Max — or a SWE-bench Verified number.
Compare every model's knowledge cutoff, max output, and context window.
Released August 26, 2026. The open-weight release is Qwen3.8-Flash-Next, a preview of the Qwen4 architecture: 125B mixture-of-experts with 6B active per token, a 51B n-gram embedding table and a 4B multi-token prediction layer. Qwen 3.8 Flash is the production API version on Qwen Cloud at $0.16/$0.47.
SWE-bench Pro 62.5 — competitive with models several times its price
Only 6B active parameters per token from a 125B mixture-of-experts, so throughput is high and hosting is cheap
991K context window at $0.16/$0.47
No published SWE-bench Verified score, only SWE-bench Pro
An architecture preview rather than a settled flagship — Qwen 3.8 Max remains Alibaba's top-end model
What people actually use Qwen 3.8 Flash for.
Volume coding work where SWE-bench Pro 62.5 is enough and cost per token dominates
Near-1M-context document processing at budget-tier rates
Self-hosted inference on modest hardware thanks to 6B active parameters per token
The nearest models people weigh against it, and what actually separates them.
vs DeepSeek V4-Flash — Against DeepSeek V4-Flash (DeepSeek), Qwen 3.8 Flash costs about 33% more per token, gives up 1x on context and answers faster. DeepSeek V4-Flash is the one to check first if the price difference matters more than the ceiling.
vs DeepSeek V4-Pro — Against DeepSeek V4-Pro (DeepSeek), Qwen 3.8 Flash runs about 52% cheaper per token, gives up 1x on context and answers faster. Take Qwen 3.8 Flash unless you specifically need what DeepSeek V4-Pro does better.
vs Devstral Small 1.1 — Against Devstral Small 1.1 (Mistral), Qwen 3.8 Flash costs about 37% more per token, takes 7.6x the context and answers faster. Devstral Small 1.1 is the one to check first if the price difference matters more than the ceiling.
Price History
→0% since Sep 1
8 data points · tracked daily since Sep 1, 2026
Cheap high-throughput coding and reasoning. Start free — no card required.
Recommendations are made independently based on real-world use and public benchmarks. See our disclosures for details.
Similar models worth checking before you commit.
A 284B-parameter (13B active) MoE workhorse re-post-trained for agentic and coding tasks — beats the V4-Pro preview on every published agent benchmark at ultra-commodity pricing.
DeepSeek's 1.6T-parameter (49B active) MoE flagship with hybrid sparse attention — near-frontier coding and reasoning at roughly a tenth of closed-rival pricing, MIT-licensed open weights.
Devstral Small 1.1 is Mistral's code-specialized small model, purpose-built for software engineering tasks including code generation, debugging, and repository-level reasoning. It succeeds Devstral Small 1.0 with improved instruction following and agentic coding capabilities at a fraction of flagship model costs.
Qwen 3.8 Flash costs $0.16 per million input tokens and $0.47 per million output tokens on the API, with cached input at $0.016 per million. A month of 10M input and 2M output tokens runs about $2.54 at list price, before any batch or caching discounts.
Qwen 3.8 Flash has a 991k tokens context window, with up to 128k tokens of output per response. That is the total of prompt plus response the model can hold in one request.
Qwen 3.8 Flash is best for cheap high-throughput coding and reasoning. It is a strong fit when that workflow matters more than the tradeoffs around budget pricing and very fast speed.
You need Alibaba's maximum capability — that is Qwen 3.8 Max — or a SWE-bench Verified number.
DeepSeek V4-Flash (DeepSeek) at $0.14/1M/1M input against Qwen 3.8 Flash's $0.16/1M/1M — roughly 33% less per token all in. Best agentic capability per dollar in the directory. Compare it first if Qwen 3.8 Flash's pricing is the thing stopping you.
DeepSeek V4-Pro — balanced against Qwen 3.8 Flash's very fast, with 1M tokens of context. Worth the swap when response time is what your users notice rather than the last few points of reasoning depth.
Newsletter
We track pricing daily. When this model drops or spikes, you'll know first.
No spam. Useful updates only. Affiliate disclosures always clearly labeled.
No reviews yet — be the first.