Amazon Polly
Audio Available Full comparison ↗ amazon/polly · by Amazon
· neural-tts
Pricing — 1 offering(s)
Standard voices
- $4.00 / 1M characters
Current
2016-11-01 → present
SCO-608 re-verification: unchanged at $4.00/1M characters.
Neural voices (NTTS)
- $16.00 / 1M characters
Current
2019-11-01 → present
SCO-608 re-verification: unchanged at $16.00/1M characters.
Long-form voices
- $100.00 / 1M characters
Current
2023-11-01 → present
SCO-608 re-verification: unchanged at $100.00/1M characters. Flag for a follow-up, not fixed here (out of…
SCO-608 re-verification: unchanged at $100.00/1M characters. Flag for a follow-up, not fixed here (out of this ticket's scope): AWS now also lists a fourth "Generative voices" tier at $30/1M characters on the live pricing page, not currently tracked as a tier on this entry — a coverage gap, not a staleness issue.
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Amazon Polly fits into a cost-aware routing setup
See how →Capability profile
How Amazon Polly rates across core capability dimensions, with the task-level evidence behind each rating.
Neural voices are natural and pleasant; not best-in-class compared to ElevenLabs or Azure Neural, but solid for most app use cases. FILLED 2026-08-19 (SCO-462 sweep, was "no independent MOS published"): the Artificial Analysis Text to Speech Arena (artificialanalysis.ai/text-to-speech/leaderboard/provider-voice) ranks "Amazon Polly Neural" #94 of 98 voices tracked (Elo 887) and "Amazon Polly Standard" dead last, #98 (Elo 817) — confirms the "not best-in-class" framing, and more starkly than the vague "community comparisons" language previously suggested.
30+ languages, 60+ voices across Standard and Neural tiers. Coverage is narrower than Azure (140+) or Google (40+) but sufficient for major world languages.
60+ voices — fewer than Azure or Google but covers major languages. No voice cloning.
Real-time streaming synthesis supported via the SynthesizeSpeech Streaming API. First-byte latency is competitive for AWS workloads.
No voice cloning or custom voice creation. Preset voices only.
Benchmarks
| Benchmark | Score | Config | Source |
|---|---|---|---|
| MOS (naturalness) | — | No independently published MOS for Polly Neural. Community comparisons generally rate it below Azure and Google Neural for naturalness, but suitable for most production use cases. | — |
Operator guidance
Best for teams running AWS-native infrastructure. At $16/1M chars (Neural) it is price-identical to Azure and Google Neural2. The 12-month free tier (5M chars/month) is the most generous among major cloud providers for new accounts. Long-form voices ($100/1M) are a niche premium for audiobook production. For voice cloning, choose ElevenLabs or PlayHT. For real-time conversational AI latency, choose Cartesia Sonic.
Use cases
- App TTS within AWS-native architectures (Lambda, EC2, Step Functions)
- Audiobook and podcast narration using Long-form voices
- Cost-optimised TTS at scale with the 12-month free tier
- Multi-language applications in the major world languages covered
Limitations
- Narrower language coverage than Azure (140+) and Google (40+)
- Neural voice quality is competitive but not best-in-class
- No voice cloning or custom voice training