Gemini 3.8 Live
Audio Available Full comparison ↗ google-deepmind/gemini-3.8-live · by Google DeepMind
· audio-native-transformer
Pricing — 1 offering(s)
Audio input tokens
- $3.00 / 1M tokens (input)
Current
2026-09-15 → present
Standard (paid tier) audio input rate; Google also states it as $0.005/min — the per-minute figure is…
Standard (paid tier) audio input rate; Google also states it as $0.005/min — the per-minute figure is Google's conversion of the token rate, so the native token unit is recorded. Text ($0.75/1M) and image/video ($1.00/1M or $0.002/min) input rates also apply when a session sends those modalities; only the audio rate is tracked as a tier, following gpt-audio-1-5-openai-audio's audio-focused convention (and so the cheaper text rate never becomes this speech model's headline price). effective_from = launch date (2026-09-15): Google lists gemini-3.8-live and gemini-3.8-live-extended-thinking on one shared price row with gemini-3.1-flash-live-preview; Wayback captures show that row without the 3.8 models on 2026-09-12 and with them, at unchanged rates, by 2026-09-16 13:42 UTC.
Audio output tokens
- $12.00 / 1M tokens (output)
Current
2026-09-15 → present
Standard (paid tier) audio output rate, including thinking tokens; Google also states it as $0.018/min. Text…
Standard (paid tier) audio output rate, including thinking tokens; Google also states it as $0.018/min. Text output is $4.50/1M (not tracked as a tier — see the input tier's notes). Grounding with Google Search: 5,000 free requests/month shared across Gemini 3.x, then $14 per 1,000. No Batch/Flex/Priority rates are published for the Live models (Live API only).
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Gemini 3.8 Live fits into a cost-aware routing setup
See how →Capability profile
How Gemini 3.8 Live rates across core capability dimensions, with the task-level evidence behind each rating.
No independent speech-to-speech quality score found as of 2026-09-27 (Google's model page and launch post give none).
Google positions the Live API as low-latency and supports barge-in, but publishes no time-to-first-audio figure for this model.
The Live API guide states 70 supported conversation languages (a Live API-level figure, not broken out per model).
No voice count or voice-selection detail found on the model page or Live API overview as of 2026-09-27.
No mention of custom or replicated voices for the Live models found in Google's docs — not assumed absent just because unmentioned.
Ratings are estimated — limited independent data is available for this model.
Operator guidance
Google's recommended default for Live API voice agents and the replacement for gemini-3.1-flash-live-preview. Same price as Gemini 3.8 Live Extended Thinking, so choose between them on behaviour, not cost: this one for the lowest latency, Extended Thinking when the agent must reason through multi-step problems mid-conversation.
Use cases
- Low-latency voice agents and real-time spoken dialogue (Google's stated default use)
- Voice and video experiences needing multimodal input in one streaming session
Limitations
- No independent quality or latency benchmark found as of 2026-09-27
- Structured outputs are not supported (Google's model page)
- thinking_level is not supported; omit thinking_config from the session setup (Google's migration notes)
- Token-billed across four rates (text/audio in, text/audio out) plus image/video input — a per-minute comparison against grok-voice-think-fast-2.0 relies on Google's own per-minute equivalents ($0.005/min in, $0.018/min out for audio)