DeepSeek V4.1-Flash
Language Available Full comparison ↗ deepseek/deepseek-v4.1-flash · by DeepSeek
· mixture-of-experts
Pricing — 1 offering(s)
Input tokens
- $0.30 / 1M tokens (input)
Current
2026-09-10 → present
SCO-636: peak-hour cache-miss input rate, the headline price. Cache-miss input is the tracked rate (same…
SCO-636: peak-hour cache-miss input rate, the headline price. Cache-miss input is the tracked rate (same convention as the other DeepSeek entries); the cache-hit rate is $0.006 peak / $0.003 off-peak, and this registry has no unit for it yet (SCO-641). The off-peak tier is half this rate and covers roughly 80% of the week's hours, so it's what most callers actually pay outside business hours. Effective from the 2026-09-10 release (API change log, https://api-docs.deepseek.com/updates; the new price is already in https://web.archive.org/web/20260910071252/https://api-docs.deepseek.com/quick_start/pricing/).
Output tokens
- $1.20 / 1M tokens (output)
Current
2026-09-10 → present
SCO-636. Peak-hour output rate; same sourcing basis as the input tier.
Off-peak input tokens
- $0.15 / 1M tokens (input)
Current
2026-09-10 → present
SCO-636: off-peak cache-miss input rate, half the peak rate. Not a headline price (attributes.processing is…
SCO-636: off-peak cache-miss input rate, half the peak rate. Not a headline price (attributes.processing is off-peak, so isHeadlineTier excludes it), but it covers roughly 80% of hours, so it's the rate most callers actually pay outside business hours. Peak window history since this row began (all other hours are off-peak): from 2026-09-10, 01:00-04:00 and 06:00-10:00 UTC Monday to Friday, weekends off-peak (2026-09-10 and 2026-09-17 snapshots). By 2026-09-22 Chinese public holidays are also off-peak in full (first seen in https://web.archive.org/web/20260922212416/https://api-docs.deepseek.com/quick_start/pricing/, and on the live page on 2026-09-28). Under the current rule peak is 35 of the week's 168 hours, so off-peak is about 79% of hours (more in holiday weeks).
Off-peak output tokens
- $0.60 / 1M tokens (output)
Current
2026-09-10 → present
SCO-636: off-peak output rate, half the peak rate. Same window history as the off-peak input tier.
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how DeepSeek V4.1-Flash fits into a cost-aware routing setup
See how →Capability profile
How DeepSeek V4.1-Flash rates across core capability dimensions, with the task-level evidence behind each rating.
1,000,000-token context, 384K maximum output (official pricing page).
Thinking and non-thinking modes. DeepSeek's change log reports GPQA Diamond 90.9 and HLE 36.8 (text-only subset); these are vendor-reported and not independently verified here.
Vendor-reported coding-agent results in the 2026-09-10 change log (Terminal-Bench 2.1 90.6, DeepSWE v1.1 74.2). No independent score yet; the modelglass-coding vertical doesn't carry this model.
Tool calls, the Responses API and the Anthropic-format API are all listed as supported on the official pricing page.
$0.30 / 1M cache-miss input and $1.20 / 1M output at peak, half that off-peak ($0.15 / $0.60), which covers roughly 80% of the week's hours. Cache hits are $0.006 peak / $0.003 off-peak. Official pricing page, checked 2026-09-28.
Ratings are estimated — limited independent data is available for this model.
Operator guidance
DeepSeek's current default model. Prefer it over V4-Pro unless you have your own evidence V4-Pro does better on your task: DeepSeek says V4.1-Flash surpassed V4-Pro on performance, cost and speed, but that's a vendor claim. Schedule batch-like work off-peak to pay half the headline rate.
Use cases
- Low-cost long-context and agentic work, especially when jobs can run off-peak (weekends, Chinese public holidays and outside 01:00-04:00 / 06:00-10:00 UTC on weekdays)
- Image-plus-text agent tasks on the DeepSeek API (V4-Pro has no vision support)
- Drop-in replacement for code still calling deepseek-v4-flash, which now routes here
Limitations
- All capability figures available so far are vendor-reported; no independent benchmark coverage checked yet, so ratings stay at unknown/moderate
- Architecture details (parameter counts, attention design) aren't published on the pages checked
- Peak-hour pricing is double the off-peak rate; costs depend on when requests run