DeepSeek · AI model API pricing
How much does DeepSeek V4 Flash cost?
The successor to deepseek-chat/reasoner (both deprecated July 24, 2026) — still one of the cheapest models here, though a July 2026 repricing roughly tripled its rates.
Rate card
| Input tokens | $0.44 / 1M |
| Cached input | $0.014 / 1M |
| Output tokens | $1.32 / 1M |
| Context window | — |
What that means per month
| Workload | Tokens / month | Cost / month |
|---|---|---|
| Prototype | 1M in · 0.3M out | $0.84/mo |
| Small app | 10M in · 3M out | $8.36/mo |
| Production app | 50M in · 15M out | $41.8/mo |
| High volume | 200M in · 60M out | $167/mo |
| Heavy platform | 1B in · 300M out | $836/mo |
Standard on-demand rates from DeepSeek's official API docs, last checked 2026-07-28. Cache-hit input bills at $0.014 per 1M at peak ($0.007 off-peak). The legacy deepseek-chat and deepseek-reasoner names map to this model. Figures use PEAK rates. DeepSeek bills by time of day: peak is 01:00-04:00 and 06:00-10:00 UTC, and off-peak (the other 17 hours) is half these prices. We publish peak so the figure can never understate a bill; halve it for off-peak traffic. Monthly figures exclude caching and batch discounts. Confirm live pricing before committing.
Where DeepSeek V4 Flash sits on price
At the reference month of 10M input and 3M output tokens, DeepSeek V4 Flash costs $8.36, which makes it the 9th cheapest of the 30 models tracked here. That is 4.9× the price of Qwen-Flash, the cheapest model at this mix ($1.70). Among the 9 budget-tier models it ranks 7th, behind Qwen-Flash at $1.70.
Cache economics
Cached input on DeepSeek V4 Flash bills at $0.014 per 1M — 3.2% of its own input rate, against a 10% median across the models here that publish a cached-read price. No model we track discounts cached reads more steeply. On the reference month, serving 80% of that input from cache takes the input side of the bill from $4.40 to $0.99, leaving output untouched at $3.96 — which is why a high reuse rate changes the ranking above rather than just shaving the total.
Cache reads only. Writes are billed separately by some providers and are excluded here — see the rate note above for DeepSeek's terms. Try it at your own hit rate on the AI cost calculator.
Cheaper alternatives
At a reference workload of 10M input / 3M output tokens a month, these cost less than DeepSeek V4 Flash ($8.36/mo):
-
Qwen-Plus $7.60/mo -
Gemini 3.1 Flash-Lite $7/mo -
MiniMax M3 $6.60/mo
DeepSeek V4 Flash head-to-head
DeepSeek V4 Flash price history
- July 2026 · reported DeepSeek roughly triples V4 pricing and moves to peak/off-peak billing
DeepSeek V4 Flash went from $0.14 in / $0.28 out to $0.44 in / $1.32 out per 1M at peak, and V4 Pro from $0.435 / $0.87 to $1.32 / $3.96 — output rates rose about 4.6x. DeepSeek also introduced time-of-day billing: peak is 01:00-04:00 and 06:00-10:00 UTC, with off-peak (the other 17 hours) at half price. Cache-hit input rose in step, taking V4 Pro's cached discount from roughly 0.8% of its input rate to about 3.3%. DeepSeek remains cheap by frontier-lab standards, but it is no longer an order of magnitude below the field.