← All models

DeepSeek V4-Flash

Language Available Full comparison ↗

deepseek/deepseek-v4-flash · by DeepSeek · mixture-of-experts

Pricing — 1 offering(s)

Input tokens

  • $0.14 / 1M tokens (input) Current 2026-08-04 → present

Output tokens

  • $0.28 / 1M tokens (output) Current 2026-08-04 → present

Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.

See how DeepSeek V4-Flash fits into a cost-aware routing setup

See how →

Capability profile

context window strong
reasoning moderate
coding moderate
tool use strong
instruction following moderate
multilingual moderate
speed strong
cost efficiency strong

Operator guidance

The default V4 choice for cost- and latency-sensitive long-context and agentic work. Step up to V4-Pro when reasoning/coding ceiling matters more than cost. Against GLM-5.2 (the other MIT Chinese-lab flagship added in the same registry batch), V4-Flash is cheaper with the same 1M context; GLM-5.2 has more (vendor-reported) coding-benchmark signal.

Use cases

Limitations

Citations