DeepSeek V4 Flash Pricing
DeepSeek ยท DeepSeek ยท Live
DeepSeek V4 Flash costs $0.22 per million input tokens and $0.66 per million output tokens on the DeepSeek API. Cached input reads are billed at $0.007 per million tokens. It supports a 1,000,000-token context window and up to 384,000 output tokens per request. The model supports tool use. Pricing verified September 14, 2026.
DeepSeek V4 Flash API pricing
DeepSeek V4 Flash API pricing is $0.22 per million input tokens and $0.66 per million output tokens, verified September 14, 2026.
- Input (per 1M tokens)
- $0.22
- Output (per 1M tokens)
- $0.66
- Cached input (per 1M tokens)
- $0.01
- Context window
- 1,000,000 tokens
- Max output
- 384,000 tokens
Capabilities
DeepSeek V4 Flash supports tool use.
Price history
See full DeepSeek V4 Flash price historyDeepSeek V4 Flash API pricing has changed 2 times on record, most recently on August 24, 2026 (increase).
- August 24, 2026: Increase โ input $0.14 โ $0.22/M (+$0.08) โ output $0.66/M
- August 17, 2026: Increase โ input $0.14/M โ output $0.28 โ $0.66/M (+$0.38)
Frequently asked questions
Short answers to the most common questions about DeepSeek V4 Flash pricing, limits, and capabilities.
Notes
Cheaper sibling of V4 Pro โ surfaced by the weekly refresh agent May 28, 2026 (source: api-docs.deepseek.com). Cache-miss input $0.14, output $0.28 per MTok. [2026-05-30] Updated Context Window 128000 โ 1000000 and Max Output Tokens 8192 โ 384000 per DeepSeek V4 Flash spec confirmed via OpenRouter + devtk.ai. Populated Cache Read Price at $0.0028 per MTok (cache-hit) per api-docs.deepseek.com โ cache-hit price is 1/50 of cache-miss input (not 1/120 as V4 Pro). Cache-hit reduction to 1/10 of launch price took effect 2026-04-26 12:15 UTC. [2026-08-17] MAJOR PRICE EVENT โ DeepSeek moved V4 to a PEAK/OFF-PEAK schedule effective 2026-08-16 16:00 UTC. Output rose $0.28 โ $0.66 off-peak and $1.32 peak (4.7ร the old rate at peak). Peak hours 01:00โ04:00 and 06:00โ10:00 UTC; all other hours off-peak. Sources: engadget.com, aipricing.guru peak/off-peak update, deepseek.ai/pricing. Stored Output = the OFF-PEAK rate ($0.66) because off-peak covers 17 of 24 hours โ schema has no peak/off-peak dimension. INPUT AND CACHE-HIT LEFT UNCHANGED ($0.14 / $0.0028): post-cutover input and cache rates were NOT published in any source found this run. ACTION FOR KHALED: recheck input/cache next run; consider adding peak/off-peak fields. This is a price increase with a time-of-day discount, not a cut. [2026-08-24] RESOLVED the open item from 2026-08-17 โ post-cutover input and cache-hit rates are now published. Input rose $0.14 โ $0.22 and Cache Read (cache-hit) rose $0.0028 โ $0.007 per MTok, both OFF-PEAK, consistent with the stored off-peak convention. Peak equivalents are $0.44 input / $0.014 cache-hit / $1.32 output. Output confirmed unchanged at $0.66 off-peak. Corroborated across deepseek.ai/pricing, aipricing.guru "V4 Peak & Off-Peak", benchlm.ai and chat-deep.ai. Cache-hit is now ~1/31 of cache-miss input (was ~1/50). Peak/off-peak schema fields still not added. [2026-09-14] Verification attempt produced no clear result โ recheck next run. The issue is not a price move but an apparent MODEL SUPERSESSION: sources this run date the stored rates ($0.22 input / $0.007 cache-hit / $0.66 output, off-peak) to the window 2026-08-16 through 2026-09-09 only, and describe DeepSeek V4 Flash as replaced from 2026-09-10 by DEEPSEEK V4.1 FLASH at $0.15 input / $0.003 cache-hit / $0.60 output off-peak (peak $0.30 / $0.006 / $1.20). V4.1 Flash is quoted with a 1M context window and 384K max output, i.e. the same spec as this row. All prices LEFT UNCHANGED โ repricing this row to V4.1 rates would silently rename the SKU, and per the task spec new models are not added automatically. ACTION FOR KHALED โ pick one: (a) reprice and rename this row to V4.1 Flash, or (b) set this row to Deprecated and add V4.1 Flash as a new row. Until then the calculator is quoting a tier that may no longer be purchasable. Sources: deepseek.ai/pricing, devtk.ai V4 Flash model page, opslyft, benchlm.ai.
