DeepSeek V4-Flash
Language Available Full comparison ↗ deepseek/deepseek-v4-flash · by DeepSeek
· mixture-of-experts
Pricing — 1 offering(s)
Input tokens
- $0.14 / 1M tokens (input) Current 2026-08-04 → present
Output tokens
- $0.28 / 1M tokens (output) Current 2026-08-04 → present
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how DeepSeek V4-Flash fits into a cost-aware routing setup
See how →Capability profile
context window strong
reasoning moderate
coding moderate
tool use strong
instruction following moderate
multilingual moderate
speed strong
cost efficiency strong
Operator guidance
The default V4 choice for cost- and latency-sensitive long-context and agentic work. Step up to V4-Pro when reasoning/coding ceiling matters more than cost. Against GLM-5.2 (the other MIT Chinese-lab flagship added in the same registry batch), V4-Flash is cheaper with the same 1M context; GLM-5.2 has more (vendor-reported) coding-benchmark signal.
Use cases
- Very-long-context tasks (whole codebases, large document sets) at low cost — the V4 generation's core pitch
- Agentic coding and tool-use workflows via the Responses API
- High-volume production workloads where the cache-hit discount and low token price compound
- Self-hosted deployment under MIT for data-control requirements
Limitations
- DeepSeek did not publish a V4-Flash-specific benchmark table; capability ratings lean on independent coverage and the V4 tech report's general claims, not verified Flash-tier numbers
- Architecture details (active-parameter count, Compressed Sparse Attention) are from independent technical coverage, not DeepSeek's own docs
- GA-date and snapshot naming vary across sources by a few weeks (V4-Flash-0731 etc.)
- An announced peak/off-peak pricing policy is not yet live and not reflected in the registry
Citations
Compare models
ComparingDeepSeek V4-Flashwith
Pick a model above to see the comparison.