← All models

Gemini 3.8 Flash

Language Available Full comparison ↗

google-deepmind/gemini-3.8-flash · by Google DeepMind · decoder-only-transformer

Pricing — 1 offering(s)

Input tokens

  • $0.75 / 1M tokens (input) Current 2026-09-02 → 2026-12-31
    Introductory rate through Dec 31 2026; scheduled step-up to $1.50/1M on Jan 1 2027 per Google's published…

    Introductory rate through Dec 31 2026; scheduled step-up to $1.50/1M on Jan 1 2027 per Google's published pricing page — not yet in effect, no forward-dated entry added ahead of it.

  • $1.50 / 1M tokens (input) Future 2027-01-01
    SCO-666: scheduled price, published in advance. Google's pricing page lists Gemini 3.8 Flash at $0.75 input /…

    SCO-666: scheduled price, published in advance. Google's pricing page lists Gemini 3.8 Flash at $0.75 input / $3.75 output "through December 31, 2026" and $1.50 / $7.50 "starting January 1, 2027". Future-dated rows aren't current until they start, so the headline stays on today's price until then.

Output tokens

  • $3.75 / 1M tokens (output) Current 2026-09-02 → 2026-12-31
    Includes thinking tokens. Introductory rate through Dec 31 2026; scheduled step-up to $7.50/1M on Jan 1 2027…

    Includes thinking tokens. Introductory rate through Dec 31 2026; scheduled step-up to $7.50/1M on Jan 1 2027 per Google's published pricing page — not yet in effect, no forward-dated entry added ahead of it.

  • $7.50 / 1M tokens (output) Future 2027-01-01
    SCO-666: scheduled price, published in advance. Google's pricing page lists Gemini 3.8 Flash at $0.75 input /…

    SCO-666: scheduled price, published in advance. Google's pricing page lists Gemini 3.8 Flash at $0.75 input / $3.75 output "through December 31, 2026" and $1.50 / $7.50 "starting January 1, 2027". Future-dated rows aren't current until they start, so the headline stays on today's price until then.

Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.

See how Gemini 3.8 Flash fits into a cost-aware routing setup

See how →

Capability profile

How Gemini 3.8 Flash rates across core capability dimensions, with the task-level evidence behind each rating.

reasoning Strong

HLE-Verified (expert-level multi-step reasoning across STEM, humanities, and professional fields): 54.9%. Vendor-reported (Google DeepMind model card) — narrowly ahead of Claude Opus 5's reported 54.4% on the same benchmark, notable for a Flash-tier model.

coding Strong

Terminal-Bench 2.1 (agentic coding tasks executed in a real terminal environment): 89.4% — fails roughly 1 task in 10. DeepSWE v1.1 (long-horizon software engineering): 73.7%. Both vendor-reported (Google DeepMind model card, deepmind.google/models/model-cards/gemini-3-8-flash/). No SWE-bench Verified/Pro score published by Google for this release; no companion modelglass-coding entry exists yet for this model as of this PR.

tool use Strong

Gemini API docs (ai.google.dev/gemini-api/docs/models/gemini-3.8-flash): function calling, code execution, and file search documented as supported; computer use listed as supported (Preview).

instruction following Strong
context window Strong

1,048,576-token input context, 65,536-token output limit.

multilingual Strong
speed Strong
cost efficiency Strong

$0.75/$3.75 per 1M input/output tokens (introductory, through 2026-12-31) — cheaper on both ends than same-tier competitor GPT-5.6 Luna ($1.00/$6.00).

Operator guidance

Google's current recommended default Flash-tier model as of 2026-09-02, superseding Gemini 3.5 Flash in this registry (Gemini 3.6/3.7 Flash were not separately tracked here). Step up to Gemini 3.1 Pro for the hardest reasoning/coding tasks; otherwise this is the cost/latency option in the current Gemini 3.x line.

Use cases

Limitations

Citations