Gemini 3.8 Live Extended Thinking
Audio Available Full comparison ↗ google-deepmind/gemini-3.8-live-extended-thinking · by Google DeepMind
· audio-native-transformer
Pricing — 1 offering(s)
Audio input tokens
- $3.00 / 1M tokens (input)
Current
2026-09-15 → present
Standard (paid tier) audio input rate; Google also states it as $0.005/min — the per-minute figure is…
Standard (paid tier) audio input rate; Google also states it as $0.005/min — the per-minute figure is Google's conversion of the token rate, so the native token unit is recorded. Text ($0.75/1M) and image/video ($1.00/1M or $0.002/min) input rates also apply when a session sends those modalities; only the audio rate is tracked as a tier, following gpt-audio-1-5-openai-audio's audio-focused convention (and so the cheaper text rate never becomes this speech model's headline price). effective_from = launch date (2026-09-15): Google lists gemini-3.8-live and gemini-3.8-live-extended-thinking on one shared price row with gemini-3.1-flash-live-preview; Wayback captures show that row without the 3.8 models on 2026-09-12 and with them, at unchanged rates, by 2026-09-16 13:42 UTC.
Audio output tokens
- $12.00 / 1M tokens (output)
Current
2026-09-15 → present
Standard (paid tier) audio output rate, including thinking tokens; Google also states it as $0.018/min. Text…
Standard (paid tier) audio output rate, including thinking tokens; Google also states it as $0.018/min. Text output is $4.50/1M (not tracked as a tier — see the input tier's notes). Grounding with Google Search: 5,000 free requests/month shared across Gemini 3.x, then $14 per 1,000. No Batch/Flex/Priority rates are published for the Live models (Live API only).
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Gemini 3.8 Live Extended Thinking fits into a cost-aware routing setup
See how →Capability profile
How Gemini 3.8 Live Extended Thinking rates across core capability dimensions, with the task-level evidence behind each rating.
No independent speech-to-speech quality score found as of 2026-09-27 (Google's model page and launch post give none).
Google positions the Live API as low-latency and supports barge-in, but publishes no time-to-first-audio figure for this model.
The Live API guide states 70 supported conversation languages (a Live API-level figure, not broken out per model).
No voice count or voice-selection detail found on the model page or Live API overview as of 2026-09-27.
No mention of custom or replicated voices for the Live models found in Google's docs — not assumed absent just because unmentioned.
Ratings are estimated — limited independent data is available for this model.
Operator guidance
Pick over plain Gemini 3.8 Live when background reasoning quality matters more than response latency. Same price row, but thinking tokens bill at the output rate, so background reasoning adds output-token cost on top of the spoken audio.
Use cases
- Real-time voice agents that must solve complex, multi-step problems mid-conversation
- Voice sessions with long-running asynchronous tool calls
Limitations
- No independent quality or latency benchmark found as of 2026-09-27
- Structured outputs are not supported (Google's model page)
- Only asynchronous, non-blocking function calling is supported; synchronous blocking mode returns a hard error (Google's model page)
- Token-billed across four rates (text/audio in, text/audio out) plus image/video input — a per-minute comparison against grok-voice-think-fast-2.0 relies on Google's own per-minute equivalents ($0.005/min in, $0.018/min out for audio)