Gemini 3.8 Flash-Lite TTS
Audio Available Full comparison ↗ google-deepmind/gemini-3.8-flash-lite-tts · by Google DeepMind
· audio-native-transformer
Pricing — 1 offering(s)
Text input tokens
- $0.50 / 1M tokens (input)
Current
2026-09-23 → present
Standard (paid tier) text input rate. Introductory rate through 2026-12-31; Google's pricing page lists a…
Standard (paid tier) text input rate. Introductory rate through 2026-12-31; Google's pricing page lists a step-up to $1.00/1M from 2027-01-01 — not yet in effect, so no forward-dated row is added ahead of it (same convention as modelglass-llm's gemini-3-8-flash-google-deepmind). effective_from = launch date (2026-09-23): the Wayback Machine's pricing-page captures show the model absent on 2026-09-22 17:49 UTC and listed at this rate by 2026-09-24 05:47 UTC.
Audio output tokens
- $6.00 / 1M tokens (output)
Current
2026-09-23 → present
Standard (paid tier) audio output rate; Google states it as equivalent to $0.0015 per 10s of audio (25 audio…
Standard (paid tier) audio output rate; Google states it as equivalent to $0.0015 per 10s of audio (25 audio tokens per second). Introductory rate through 2026-12-31; steps up to $12.00/1M ($0.003 per 10s) from 2027-01-01 — not yet in effect, no forward-dated row.
Batch text input tokens
- $0.25 / 1M tokens (input)
Current
2026-09-23 → present
Batch API rate (secondary tier per the SCO-640 headline rule — never the list price). Introductory through…
Batch API rate (secondary tier per the SCO-640 headline rule — never the list price). Introductory through 2026-12-31; $0.50/1M from 2027-01-01, not yet in effect.
Batch audio output tokens
- $3.00 / 1M tokens (output)
Current
2026-09-23 → present
Batch API rate (secondary tier per the SCO-640 headline rule). Introductory through 2026-12-31; $6.00/1M from…
Batch API rate (secondary tier per the SCO-640 headline rule). Introductory through 2026-12-31; $6.00/1M from 2027-01-01, not yet in effect.
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Gemini 3.8 Flash-Lite TTS fits into a cost-aware routing setup
See how →Capability profile
How Gemini 3.8 Flash-Lite TTS rates across core capability dimensions, with the task-level evidence behind each rating.
Google positions it for high throughput, low latency and cost efficiency rather than maximum fidelity. No independent naturalness score found as of 2026-09-27; not rated on vendor positioning alone.
30 prebuilt studio voices, plus an Extended Voice Library of hundreds more (GET /v1beta/voices) and Voice design personas generated from a text description (Google's speech-generation guide).
Voice replication from reference plus consent audio, stored persistently (up to 200 custom voices per project, 1-year retention) or as stateless client-held keys (Google's speech-generation guide).
101 languages per Google's model page, with automatic input-language detection.
Streaming supported (headerless 16-bit PCM at 24 kHz); Google publishes no time-to-first-audio figure.
Ratings are estimated — limited independent data is available for this model.
Operator guidance
Default Gemini TTS pick for cost-sensitive, high-volume or latency-sensitive speech ($6 vs $9 per 1M audio output tokens against Gemini 3.8 Flash TTS, introductory rates). Switching to Flash TTS is a one-parameter change, so step up only for narration, multi-speaker acting or the extra dialect coverage.
Use cases
- High-volume production speech generation
- Real-time voice-agent cascades (LLM → TTS)
- Read-aloud features and everyday single-speaker generation
Limitations
- No independent quality or latency benchmark found as of 2026-09-27 — capability ratings above rest on Google's own documentation
- Pricing is introductory through 2026-12-31 and doubles on 2027-01-01 (Google's pricing page)
- Token-billed: comparing against per-character TTS needs a tokens-per-character ratio Google does not publish (25 audio tokens per second of output is published)
- Inline text directions from earlier Gemini TTS prompts may now be spoken aloud — prompts need migrating to speech_metadata