← All models

DeepSeek V4.1-Flash

Language Available Full comparison ↗

deepseek/deepseek-v4.1-flash · by DeepSeek · mixture-of-experts

Pricing — 1 offering(s)

Input tokens

  • $0.30 / 1M tokens (input) Current 2026-09-10 → present
    SCO-636: peak-hour cache-miss input rate, the headline price. Cache-miss input is the tracked rate (same…

    SCO-636: peak-hour cache-miss input rate, the headline price. Cache-miss input is the tracked rate (same convention as the other DeepSeek entries); the cache-hit rate is $0.006 peak / $0.003 off-peak, and this registry has no unit for it yet (SCO-641). The off-peak tier is half this rate and covers roughly 80% of the week's hours, so it's what most callers actually pay outside business hours. Effective from the 2026-09-10 release (API change log, https://api-docs.deepseek.com/updates; the new price is already in https://web.archive.org/web/20260910071252/https://api-docs.deepseek.com/quick_start/pricing/).

Output tokens

  • $1.20 / 1M tokens (output) Current 2026-09-10 → present

    SCO-636. Peak-hour output rate; same sourcing basis as the input tier.

Off-peak input tokens

  • $0.15 / 1M tokens (input) Current 2026-09-10 → present
    SCO-636: off-peak cache-miss input rate, half the peak rate. Not a headline price (attributes.processing is…

    SCO-636: off-peak cache-miss input rate, half the peak rate. Not a headline price (attributes.processing is off-peak, so isHeadlineTier excludes it), but it covers roughly 80% of hours, so it's the rate most callers actually pay outside business hours. Peak window history since this row began (all other hours are off-peak): from 2026-09-10, 01:00-04:00 and 06:00-10:00 UTC Monday to Friday, weekends off-peak (2026-09-10 and 2026-09-17 snapshots). By 2026-09-22 Chinese public holidays are also off-peak in full (first seen in https://web.archive.org/web/20260922212416/https://api-docs.deepseek.com/quick_start/pricing/, and on the live page on 2026-09-28). Under the current rule peak is 35 of the week's 168 hours, so off-peak is about 79% of hours (more in holiday weeks).

Off-peak output tokens

  • $0.60 / 1M tokens (output) Current 2026-09-10 → present

    SCO-636: off-peak output rate, half the peak rate. Same window history as the off-peak input tier.

Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.

See how DeepSeek V4.1-Flash fits into a cost-aware routing setup

See how →

Capability profile

How DeepSeek V4.1-Flash rates across core capability dimensions, with the task-level evidence behind each rating.

context window Strong

1,000,000-token context, 384K maximum output (official pricing page).

reasoning Unknown

Thinking and non-thinking modes. DeepSeek's change log reports GPQA Diamond 90.9 and HLE 36.8 (text-only subset); these are vendor-reported and not independently verified here.

coding Unknown

Vendor-reported coding-agent results in the 2026-09-10 change log (Terminal-Bench 2.1 90.6, DeepSWE v1.1 74.2). No independent score yet; the modelglass-coding vertical doesn't carry this model.

tool use Moderate

Tool calls, the Responses API and the Anthropic-format API are all listed as supported on the official pricing page.

cost efficiency Strong

$0.30 / 1M cache-miss input and $1.20 / 1M output at peak, half that off-peak ($0.15 / $0.60), which covers roughly 80% of the week's hours. Cache hits are $0.006 peak / $0.003 off-peak. Official pricing page, checked 2026-09-28.

Ratings are estimated — limited independent data is available for this model.

Operator guidance

DeepSeek's current default model. Prefer it over V4-Pro unless you have your own evidence V4-Pro does better on your task: DeepSeek says V4.1-Flash surpassed V4-Pro on performance, cost and speed, but that's a vendor claim. Schedule batch-like work off-peak to pay half the headline rate.

Use cases

Limitations

Citations