| Metric | Value |
|---|---|
| Input price | $0.14 / 1M tokens |
| Output price | $0.28 / 1M tokens |
| Context window | 1M |
| Best for | Cheapest frontier |
Example monthly cost
At a sample workload of 50,000 requests/month with 8,000 context tokens, 1,000 new input tokens, and 500 output tokens per request, DeepSeek V4 Flash costs roughly $70.00/month (uncached). Turn on prompt caching and the context portion drops to about 10%, often cutting the bill by half. Estimate your own numbers with the AI Cost Calculator.
How DeepSeek V4 Flash compares
DeepSeek V4 Flash sits among the current frontier and open-weight models. To see every model side by side — context window, input/output price, and best-fit use case — open the LLM Model Compare tool and filter by what you need (coding, long context, cheap volume, or reasoning).
Frequently asked questions
- How much does DeepSeek V4 Flash cost per 1M tokens?
- Input tokens are $0.14 per 1M and output tokens are $0.28 per 1M (indicative 2026 rates). Output costs more because generating tokens uses more compute.
- What is the DeepSeek V4 Flash context window?
- 1M. A larger context window lets you send more code, docs, or conversation history per request without chunking.
- Is this DeepSeek V4 Flash pricing current?
- Rates are a snapshot captured at build time. AI pricing moves often and many providers offer cached-token discounts, so confirm on the official pricing page before committing a budget.
- How do I estimate my own DeepSeek V4 Flash bill?
- Use the AI Cost Calculator — enter your tokens and volume and see the monthly cost instantly, with and without prompt caching.