DeepSeek V4-Pro
DeepSeek's 1.6T-parameter (49B active) MoE flagship with hybrid sparse attention — near-frontier coding and reasoning at roughly a tenth of closed-rival pricing, MIT-licensed open weights.
DeepSeek's open, cheap new default — V4-Pro now routes to it.
Cheap coding and agent work with open weights
You need long single outputs (32K cap) or independently verified scores before you commit.
Compare every model's knowledge cutoff, max output, and context window.
Released September 10, 2026; previous V4 Flash and V4 Flash Vision endpoints taken offline the same day. From September 14, 2026, requests to deepseek-v4-pro are served by V4.1 Flash at Flash rates until a V4.1-Pro ships. Gateway id deepseek/deepseek-v4.1-flash; gateway price $0.30/$1.20 per 1M (varies by host; DeepSeek's own API uses peak and off-peak rates). 1,048,576 context, 32,768 max output on the gateway. Weights on Hugging Face, MIT licence. Verified October 10, 2026.
MIT-licensed weights: download, modify and self-host
1M context with native image input, at roughly $0.30/$1.20 per 1M on the gateway
Adjustable reasoning effort to trade accuracy against cost
DeepSeek's coding claims are self-reported and not yet independently verified; it trails on hard reasoning by DeepSeek's own account
Price varies by host and by time of day — check the rate you will actually pay
32K max output on the gateway
What people actually use DeepSeek V4.1 Flash for.
High-volume coding assistance where review catches what the model misses
Self-hosted deployments — the weights are on Hugging Face under the MIT licence
Image understanding without a separate vision model
The nearest models people weigh against it, and what actually separates them.
vs DeepSeek V4-Pro — Against DeepSeek V4-Pro (DeepSeek), DeepSeek V4.1 Flash costs about 13% more per token, takes 1x the context and answers faster. DeepSeek V4-Pro is the one to check first if the price difference matters more than the ceiling.
vs DeepSeek V4-Flash — Against DeepSeek V4-Flash (DeepSeek), DeepSeek V4.1 Flash costs about 72% more per token and takes 1x the context. DeepSeek V4-Flash is the one to check first if the price difference matters more than the ceiling.
vs Gemini 3.8 Flash — Against Gemini 3.8 Flash (Google), DeepSeek V4.1 Flash runs about 67% cheaper per token, takes 1x the context and answers slower. Which one wins depends on whether context depth or latency is your constraint.
Cheap coding and agent work with open weights. Start free — no card required.
Recommendations are made independently based on real-world use and public benchmarks. See our disclosures for details.
Similar models worth checking before you commit.
DeepSeek's 1.6T-parameter (49B active) MoE flagship with hybrid sparse attention — near-frontier coding and reasoning at roughly a tenth of closed-rival pricing, MIT-licensed open weights.
A 284B-parameter (13B active) MoE workhorse re-post-trained for agentic and coding tasks — beats the V4-Pro preview on every published agent benchmark at ultra-commodity pricing.
Google's September 2, 2026 Flash model, positioned as the agentic workhorse of the Gemini 3 family — stronger coding and terminal work than 3.7 Flash at the same $0.75/$3.75 price.
Pricing moves, ranking shifts, and capability updates.
DeepSeek released V4.1 Flash on September 10, 2026, took the previous V4 Flash endpoints offline the same day, and from September 14 began serving all V4-Pro API requests with V4.1 Flash at Flash rates until a V4.1-Pro ships. The model is natively multimodal, has a 1M-token context and adjustable reasoning effort, and its weights are on Hugging Face under the MIT licence. DeepSeek says it beats V4-Pro on coding and agent tasks; those figures are self-reported. Hosted prices vary by provider; the Vercel AI Gateway lists $0.30/$1.20 per million tokens. Verified October 10, 2026.
View modelDeepSeek V4.1 Flash costs $0.3 per million input tokens and $1.2 per million output tokens on the API. A month of 10M input and 2M output tokens runs about $5.40 at list price, before any batch or caching discounts.
DeepSeek V4.1 Flash has a 1.0M tokens context window, with up to 33k tokens of output per response. That is the total of prompt plus response the model can hold in one request.
DeepSeek V4.1 Flash is best for cheap coding and agent work with open weights. It is a strong fit when that workflow matters more than the tradeoffs around budget pricing and fast speed.
You need long single outputs (32K cap) or independently verified scores before you commit.
DeepSeek V4-Pro (DeepSeek) at $0.43/1M/1M input against DeepSeek V4.1 Flash's $0.30/1M/1M — roughly 13% less per token all in. Former DeepSeek flagship — API requests now run on V4.1 Flash. Compare it first if DeepSeek V4.1 Flash's pricing is the thing stopping you.
Gemini 3.8 Flash — very fast against DeepSeek V4.1 Flash's fast, with 1M tokens of context. Worth the swap when response time is what your users notice rather than the last few points of reasoning depth.
Newsletter
We track pricing daily. When this model drops or spikes, you'll know first.
No spam. Useful updates only. Affiliate disclosures always clearly labeled.
No reviews yet — be the first.