← All models

Google Cloud TTS

Audio Available Full comparison ↗

google-cloud/tts · by Google · neural-tts

Pricing — 1 offering(s)

Standard voices

  • $4.00 / 1M characters Current 2018-03-01 → present

    SCO-608 re-verification: unchanged at $4.00/1M characters.

WaveNet / Neural2 voices (superseded — see wavenet-voices / neural2-voices below)

No current price · last price $16.00 / 1M characters, until 2026-09-19

  • $16.00 / 1M characters Historical 2018-03-01 → 2026-09-19
    SCO-608 (2026-09-20): split into two separate tiers below — WaveNet and Neural2 no longer share one price on…

    SCO-608 (2026-09-20): split into two separate tiers below — WaveNet and Neural2 no longer share one price on Google's current pricing page. This combined tier is kept (capped, not deleted) as the accurate historical record of when they did.

WaveNet voices

  • $4.00 / 1M characters Current 2026-09-20 → present
    SCO-608: new tier, split out of wavenet-neural2-voices. Google's current "Legacy TTS models" pricing table…

    SCO-608: new tier, split out of wavenet-neural2-voices. Google's current "Legacy TTS models" pricing table lists WaveNet at the same $4/1M rate as Standard voices — a real repricing down from the $16/1M it shared with Neural2 before. Free tier: 0-4M characters/month (same band as Standard).

Neural2 voices

  • $16.00 / 1M characters Current 2026-09-20 → present
    SCO-608: new tier, split out of wavenet-neural2-voices. Neural2 stayed at $16/1M characters — unchanged, just…

    SCO-608: new tier, split out of wavenet-neural2-voices. Neural2 stayed at $16/1M characters — unchanged, just no longer bundled with WaveNet's now-lower rate. Free tier: 0-1M characters/month.

Studio voices

  • $160.00 / 1M characters Current 2022-01-01 → present

    SCO-608 re-verification: unchanged at $160.00/1M characters.

Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.

See how Google Cloud TTS fits into a cost-aware routing setup

See how →

About Google Cloud TTS

Google Cloud TTS is not one model but three generations under one API. WaveNet, from 2016, was Google DeepMind's original neural TTS architecture (the source paper reported a 4.21 MOS on English). Neural2, from 2022, is a successor with improved naturalness and prosody. Studio voices, also 2022, are human-supervised and the highest-quality tier. All run on Google's TPU infrastructure, covering 40+ languages and 380+ voices, with SSML support but no voice cloning — Custom Voice is a separate contracted service.

It is a REST API oriented to batch synthesis rather than real-time conversational use. Pricing is $16 per 1M characters for WaveNet/Neural2 — identical to Azure Neural TTS and Amazon Polly Neural — with Studio voices at $160 per 1M, and a 1M-character-per-month free tier for the WaveNet and Neural2 tiers.

The SCO-462 review (2026-08-19) surfaced a finding worth keeping separate rather than smoothing into one quality claim: the Artificial Analysis Text-to-Speech Arena tracks the three tiers individually, and while Studio ranks #40 of 98 voices (Elo 1081, clearly the strongest), WaveNet (#90, Elo 913) and Neural2 (#93, Elo 890) rank surprisingly low. Google's own published evaluations rate Neural2 highly; there is no independent MOS against competitors beyond that Arena.

Capability profile

How Google Cloud TTS rates across core capability dimensions, with the task-level evidence behind each rating.

voice naturalness Strong

Neural2 and WaveNet voices are highly natural for an API TTS; Studio voices approach human quality on professional narration content. FILLED 2026-08-19 (SCO-462 sweep, was "no independent MOS published"): the Artificial Analysis Text to Speech Arena (artificialanalysis.ai/text-to-speech/leaderboard/provider-voice) tracks all three tiers separately: "Studio" #40 of 98 (Elo 1081, clearly the strongest of the three, consistent with this doc's "approach human quality" framing), "WaveNet" #90 (Elo 913), "Neural2" #93 (Elo 890) — WaveNet and Neural2 rank surprisingly low relative to Studio, a real finding worth noting rather than smoothing into one "highly natural" claim for all three.

language support Strong

40+ languages, 380+ voices across Standard, WaveNet/Neural2, and Studio tiers.

voice variety Strong

380+ voices across the portfolio; no voice cloning, but broad preset coverage.

streaming latency Moderate

REST API primarily for batch synthesis. SSML and audio profile support adds processing time; not purpose-built for real-time conversational AI.

cloning support Weak

No voice cloning. Preset voices only; Custom Voice is available via a separate contracted service.

Benchmarks

Benchmark Score Config Source
MOS (naturalness) — Neural2 voices score highly in Google's own published evaluations. No independent MOS published against competitors. —
WaveNet original MOS (published) 4.21 MOS (1–5) Original WaveNet paper MOS on English, not directly comparable to current Neural2 quality. source ↗

Operator guidance

Best for teams already on Google Cloud who want tight ecosystem integration and a large per-month free tier (1M WaveNet/Neural2 chars/month). At $16/1M chars (WaveNet/Neural2) it is price-identical to Azure and Polly Neural. Choose ElevenLabs for naturalness, Cartesia Sonic for latency- critical conversational AI, or OpenAI TTS-1 for the lowest price.

Use cases

Limitations

Citations