DeepSeek V4-Pro
DeepSeek's 1.6T-parameter (49B active) MoE flagship with hybrid sparse attention — near-frontier coding and reasoning at roughly a tenth of closed-rival pricing, MIT-licensed open weights.
Best agentic capability per dollar in the directory.
High-volume agentic coding and tool-use pipelines
You need vision input or frontier-grade reasoning on the hardest tasks.
Compare every model's knowledge cutoff, max output, and context window.
Official V4-Flash-0731 release July 31, 2026; weights on Hugging Face, API in public beta. Only DeepSeek model supporting the Responses API. DeepSeek has warned of a future price increase.
Terminal-Bench 2.1 82.7 — up from 61.8 in the April preview, beating V4-Pro (Preview) on all nine published agent benchmarks
Strong tool-calling and security-task results (Toolathlon-Verified 70.3, Cybergym 76.7)
$0.14/$0.28 per 1M with 1M context and MIT-licensed weights
Well behind GPT-5.6, Opus-class, and Gemini frontier models on the hardest reasoning and long-horizon work
Text-only, and several headline numbers come from DeepSeek's own unreleased eval framework
What people actually use DeepSeek V4-Flash for.
Agent pipelines at $0.14/1M input — Terminal-Bench 2.1 82.7 rivals models 30x its price
Tool-calling workloads (Toolathlon-Verified 70.3) with 2,500 concurrent requests
Self-hosting in ~110 GB at 3-bit quantization under MIT license
The nearest models people weigh against it, and what actually separates them.
vs DeepSeek V4-Pro — Against DeepSeek V4-Pro (DeepSeek), DeepSeek V4-Flash runs about 68% cheaper per token and answers faster. Take DeepSeek V4-Flash unless you specifically need what DeepSeek V4-Pro does better.
vs Devstral Small 1.1 — Against Devstral Small 1.1 (Mistral), DeepSeek V4-Flash costs about 5% more per token and takes 7.6x the context. Devstral Small 1.1 is the one to check first if the price difference matters more than the ceiling.
vs GLM-5.2 — Against GLM-5.2 (Z.ai), DeepSeek V4-Flash runs about 93% cheaper per token and answers faster. Take DeepSeek V4-Flash unless you specifically need what GLM-5.2 does better.
Price History
→0% since Aug 7
26 data points · tracked daily since Aug 7, 2026
High-volume agentic coding and tool-use pipelines. Start free — no card required.
Recommendations are made independently based on real-world use and public benchmarks. See our disclosures for details.
Similar models worth checking before you commit.
DeepSeek's 1.6T-parameter (49B active) MoE flagship with hybrid sparse attention — near-frontier coding and reasoning at roughly a tenth of closed-rival pricing, MIT-licensed open weights.
Devstral Small 1.1 is Mistral's code-specialized small model, purpose-built for software engineering tasks including code generation, debugging, and repository-level reasoning. It succeeds Devstral Small 1.0 with improved instruction following and agentic coding capabilities at a fraction of flagship model costs.
Z.ai's MIT-licensed open-weight flagship — the top open-weights coding model of mid-2026, beating GPT-5.5 on agentic coding benchmarks at roughly a sixth of the cost.
DeepSeek V4-Flash costs $0.14 per million input tokens and $0.28 per million output tokens on the API. A month of 10M input and 2M output tokens runs about $1.96 at list price, before any batch or caching discounts.
DeepSeek V4-Flash has a 1M tokens context window, with up to 384k tokens of output per response. That is the total of prompt plus response the model can hold in one request.
DeepSeek V4-Flash's training data runs through May 2025, and the model was released on April 23, 2026. For anything after that date it needs web search or documents in the prompt.
DeepSeek V4-Flash is best for high-volume agentic coding and tool-use pipelines. It is a strong fit when that workflow matters more than the tradeoffs around budget pricing and fast speed.
You need vision input or frontier-grade reasoning on the hardest tasks.
Devstral Small 1.1 (Mistral) at $0.10/1M/1M input against DeepSeek V4-Flash's $0.14/1M/1M — roughly 5% less per token all in. The best dollar-for-dollar coding model for agentic pipelines that doesn't need to do anything else. Compare it first if DeepSeek V4-Flash's pricing is the thing stopping you.
DeepSeek V4-Pro — balanced against DeepSeek V4-Flash's fast, with 1M tokens of context. Worth the swap when response time is what your users notice rather than the last few points of reasoning depth.
Newsletter
We track pricing daily. When this model drops or spikes, you'll know first.
No spam. Useful updates only. Affiliate disclosures always clearly labeled.
No reviews yet — be the first.