Google Cloud TTS
Audio Available Full comparison ↗ google-cloud/tts · by Google
· neural-tts
Pricing — 1 offering(s)
Standard voices
- $4.00 / 1M characters
Current
2018-03-01 → present
SCO-608 re-verification: unchanged at $4.00/1M characters.
WaveNet / Neural2 voices (superseded — see wavenet-voices / neural2-voices below)
No current price · last price $16.00 / 1M characters, until 2026-09-19
- $16.00 / 1M characters
Historical
2018-03-01 → 2026-09-19
SCO-608 (2026-09-20): split into two separate tiers below — WaveNet and Neural2 no longer share one price on…
SCO-608 (2026-09-20): split into two separate tiers below — WaveNet and Neural2 no longer share one price on Google's current pricing page. This combined tier is kept (capped, not deleted) as the accurate historical record of when they did.
WaveNet voices
- $4.00 / 1M characters
Current
2026-09-20 → present
SCO-608: new tier, split out of wavenet-neural2-voices. Google's current "Legacy TTS models" pricing table…
SCO-608: new tier, split out of wavenet-neural2-voices. Google's current "Legacy TTS models" pricing table lists WaveNet at the same $4/1M rate as Standard voices — a real repricing down from the $16/1M it shared with Neural2 before. Free tier: 0-4M characters/month (same band as Standard).
Neural2 voices
- $16.00 / 1M characters
Current
2026-09-20 → present
SCO-608: new tier, split out of wavenet-neural2-voices. Neural2 stayed at $16/1M characters — unchanged, just…
SCO-608: new tier, split out of wavenet-neural2-voices. Neural2 stayed at $16/1M characters — unchanged, just no longer bundled with WaveNet's now-lower rate. Free tier: 0-1M characters/month.
Studio voices
- $160.00 / 1M characters
Current
2022-01-01 → present
SCO-608 re-verification: unchanged at $160.00/1M characters.
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Google Cloud TTS fits into a cost-aware routing setup
See how →About Google Cloud TTS
Google Cloud TTS is not one model but three generations under one API. WaveNet, from 2016, was Google DeepMind's original neural TTS architecture (the source paper reported a 4.21 MOS on English). Neural2, from 2022, is a successor with improved naturalness and prosody. Studio voices, also 2022, are human-supervised and the highest-quality tier. All run on Google's TPU infrastructure, covering 40+ languages and 380+ voices, with SSML support but no voice cloning — Custom Voice is a separate contracted service.
It is a REST API oriented to batch synthesis rather than real-time conversational use. Pricing is $16 per 1M characters for WaveNet/Neural2 — identical to Azure Neural TTS and Amazon Polly Neural — with Studio voices at $160 per 1M, and a 1M-character-per-month free tier for the WaveNet and Neural2 tiers.
The SCO-462 review (2026-08-19) surfaced a finding worth keeping separate rather than smoothing into one quality claim: the Artificial Analysis Text-to-Speech Arena tracks the three tiers individually, and while Studio ranks #40 of 98 voices (Elo 1081, clearly the strongest), WaveNet (#90, Elo 913) and Neural2 (#93, Elo 890) rank surprisingly low. Google's own published evaluations rate Neural2 highly; there is no independent MOS against competitors beyond that Arena.
Capability profile
How Google Cloud TTS rates across core capability dimensions, with the task-level evidence behind each rating.
Neural2 and WaveNet voices are highly natural for an API TTS; Studio voices approach human quality on professional narration content. FILLED 2026-08-19 (SCO-462 sweep, was "no independent MOS published"): the Artificial Analysis Text to Speech Arena (artificialanalysis.ai/text-to-speech/leaderboard/provider-voice) tracks all three tiers separately: "Studio" #40 of 98 (Elo 1081, clearly the strongest of the three, consistent with this doc's "approach human quality" framing), "WaveNet" #90 (Elo 913), "Neural2" #93 (Elo 890) — WaveNet and Neural2 rank surprisingly low relative to Studio, a real finding worth noting rather than smoothing into one "highly natural" claim for all three.
40+ languages, 380+ voices across Standard, WaveNet/Neural2, and Studio tiers.
380+ voices across the portfolio; no voice cloning, but broad preset coverage.
REST API primarily for batch synthesis. SSML and audio profile support adds processing time; not purpose-built for real-time conversational AI.
No voice cloning. Preset voices only; Custom Voice is available via a separate contracted service.
Benchmarks
| Benchmark | Score | Config | Source |
|---|---|---|---|
| MOS (naturalness) | — | Neural2 voices score highly in Google's own published evaluations. No independent MOS published against competitors. | — |
| WaveNet original MOS (published) | 4.21 MOS (1–5) | Original WaveNet paper MOS on English, not directly comparable to current Neural2 quality. | source ↗ |
Operator guidance
Best for teams already on Google Cloud who want tight ecosystem integration and a large per-month free tier (1M WaveNet/Neural2 chars/month). At $16/1M chars (WaveNet/Neural2) it is price-identical to Azure and Polly Neural. Choose ElevenLabs for naturalness, Cartesia Sonic for latency- critical conversational AI, or OpenAI TTS-1 for the lowest price.
Use cases
- Application TTS for users already in the Google Cloud ecosystem
- Multi-language content production across 40+ languages
- Professional narration and content with Studio voices
- SSML-heavy workflows requiring pitch, speed, and break control
Limitations
- No voice cloning (Custom Voice is a contracted enterprise offering)
- REST API; not purpose-built for real-time streaming TTS
- Studio voices at $160/1M chars are significantly more expensive than Neural2