← All models

Gemini 3.8 Flash TTS

Audio Available Full comparison ↗

google-deepmind/gemini-3.8-flash-tts · by Google DeepMind · audio-native-transformer

Pricing — 1 offering(s)

Text input tokens

  • $0.50 / 1M tokens (input) Current 2026-09-23 → present
    Standard (paid tier) text input rate. Introductory rate through 2026-12-31; Google's pricing page lists a…

    Standard (paid tier) text input rate. Introductory rate through 2026-12-31; Google's pricing page lists a step-up to $1.00/1M from 2027-01-01 — not yet in effect, so no forward-dated row is added ahead of it (same convention as modelglass-llm's gemini-3-8-flash-google-deepmind). effective_from = launch date (2026-09-23): the Wayback Machine's pricing-page captures show the model absent on 2026-09-22 17:49 UTC and listed at this rate by 2026-09-24 05:47 UTC.

Audio output tokens

  • $9.00 / 1M tokens (output) Current 2026-09-23 → present
    Standard (paid tier) audio output rate; Google states it as equivalent to $0.00225 per 10s of audio (25 audio…

    Standard (paid tier) audio output rate; Google states it as equivalent to $0.00225 per 10s of audio (25 audio tokens per second). Introductory rate through 2026-12-31; steps up to $18.00/1M ($0.0045 per 10s) from 2027-01-01 — not yet in effect, no forward-dated row.

Batch text input tokens

  • $0.25 / 1M tokens (input) Current 2026-09-23 → present
    Batch API rate (secondary tier per the SCO-640 headline rule — never the list price). Introductory through…

    Batch API rate (secondary tier per the SCO-640 headline rule — never the list price). Introductory through 2026-12-31; $0.50/1M from 2027-01-01, not yet in effect.

Batch audio output tokens

  • $4.50 / 1M tokens (output) Current 2026-09-23 → present
    Batch API rate (secondary tier per the SCO-640 headline rule). Introductory through 2026-12-31; $9.00/1M from…

    Batch API rate (secondary tier per the SCO-640 headline rule). Introductory through 2026-12-31; $9.00/1M from 2027-01-01, not yet in effect.

Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.

See how Gemini 3.8 Flash TTS fits into a cost-aware routing setup

See how →

Capability profile

How Gemini 3.8 Flash TTS rates across core capability dimensions, with the task-level evidence behind each rating.

voice naturalness Unknown

Google positions it as its flagship creative TTS tier (studio-grade fidelity, expressive acting, long-form multi-turn stability). No independent naturalness score found as of 2026-09-27; not rated on vendor positioning alone.

voice variety Strong

30 prebuilt studio voices, plus an Extended Voice Library of hundreds more (GET /v1beta/voices) and Voice design personas generated from a text description (Google's speech-generation guide).

cloning support Strong

Voice replication from reference plus consent audio, stored persistently (up to 200 custom voices per project, 1-year retention) or as stateless client-held keys (Google's speech-generation guide).

language support Strong

130 languages per Google's model page, with automatic input-language detection.

streaming latency Unknown

Streaming supported (headerless 16-bit PCM at 24 kHz); Google publishes no time-to-first-audio figure.

Ratings are estimated — limited independent data is available for this model.

Operator guidance

Choose over Gemini 3.8 Flash-Lite TTS when fidelity, acting nuance or dialect coverage matter more than cost: same API schema and voices, 130 vs 101 languages, $9 vs $6 per 1M audio output tokens (introductory rates). For high-volume single-speaker read-aloud or voice-agent cascades, Google itself recommends Flash-Lite.

Use cases

Limitations

Citations