Grok Voice Think Fast 2.0
Audio Available Full comparison ↗ xai/grok-voice-think-fast-2.0 · by xAI
· audio-native-transformer
Pricing — 1 offering(s)
Speech-to-speech audio
- $0.080 / minute Current 2026-07-29 → present
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Grok Voice Think Fast 2.0 fits into a cost-aware routing setup
See how →Capability profile
Operator guidance
Default choice among xAI's own Grok Voice line for new integrations — it superseded grok-voice-think-fast-1.0 as the grok-voice-latest alias on 2026-08-05 (1.0 remains available by pinning the version explicitly). Choose this specifically (over pinning 1.0) for the latency and reasoning-efficiency gains xAI cites: 0.70s time-to-first-audio and lower reasoning-token overhead. Cross-vendor, this registry has no other speech-to-speech peer priced the same way (gpt-audio-1.5 is token-billed; flash-v2-5 is TTS-only, no audio input) — pricing model, not just quality, is the deciding factor if a caller is already committed to per-minute billing.
Use cases
- Real-time conversational voice agents where response latency and in-conversation reasoning both matter (xAI's own target scenario: tool calls that execute mid-response rather than after a pause)
- Noisy or telephony-quality audio environments — xAI's own post frames the model's design focus as "real-world settings, with substantial background noise and telephony compression," and cites large transcription-accuracy gains in noisy conditions specifically (see limitations for the exact figures and their caveats)
- Multilingual voice agents needing broad (24-language-evaluated) coverage without per-language quality being individually verified
Limitations
- Quality/latency comparisons against GPT-Realtime-2.1 and Gemini 3.1 Flash are all sourced from xAI's own launch post — independently confirming they used identical test conditions was not possible from that source alone (the Speech-to-Speech Quality Index sub-score is attributed to Artificial Analysis, but the GPT/Gemini comparison figures alongside it are presented by xAI, not independently re-verified here)
- Transcription-accuracy claims ("1.5-2.0x improvement relative to Deepgram Nova 3 and ElevenLabs Scribe v2," "~10x improvement... in noisy environments") are relative multipliers, not absolute accuracy figures — the baseline error rates they're multiplying against aren't given in the same post
- No specific list of the 24 evaluated languages published, and no per-language quality breakdown — see language-support note above
- No information found on voice cloning, custom voice creation, or voice/persona variety in either direction
- No published parameter count or detailed architecture beyond xAI's own high-level framing