OpenAI TTS-1 HD
Audio Available Full comparison ↗ openai/tts-1-hd · by OpenAI
· neural-tts
Pricing — 1 offering(s)
Text characters
- $30.00 / 1M characters
Current
2023-11-06 → present
SCO-608 re-verification: unchanged at $30/1M characters. openai.com/api/pricing/ 403s to automated fetches…
SCO-608 re-verification: unchanged at $30/1M characters. openai.com/api/pricing/ 403s to automated fetches (Cloudflare bot check, not dead); confirmed instead via developers.openai.com/api/docs/pricing, which lists the same figure. SCO-617: citation URL itself updated to match — the old URL now redirects to an unrelated ChatGPT Business seat-pricing page.
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how OpenAI TTS-1 HD fits into a cost-aware routing setup
See how →Capability profile
How OpenAI TTS-1 HD rates across core capability dimensions, with the task-level evidence behind each rating.
Noticeably higher audio fidelity than TTS-1. Better prosody and less metallic quality. FILLED 2026-08-19 (SCO-462 sweep, was "no independently published MOS"): the Artificial Analysis Text to Speech Arena (artificialanalysis.ai/text-to-speech/leaderboard/provider-voice) ranks "TTS-1 HD" #30 of 98 voices tracked, Elo 1108 — mid-table, ahead of TTS-1 itself (#34, Elo 1094; see tts-1.yaml).
Same 57-language support as TTS-1.
Same 6 preset voices as TTS-1; no cloning or custom voices.
Slightly higher latency than TTS-1 due to higher fidelity processing.
No voice cloning, no fine-tuning, no SSML. Same limitations as TTS-1.
Benchmarks
| Benchmark | Score | Config | Source |
|---|---|---|---|
| MOS (naturalness) | — | No independently published MOS. Vendor-described as higher fidelity than TTS-1; community rates it between TTS-1 and ElevenLabs Multilingual v2. | — |
Operator guidance
Choose TTS-1 HD when output quality is more important than cost — the 2× price premium is justified for final production audio. For real-time apps and volume workloads, TTS-1 is sufficient. ElevenLabs Multilingual v2 exceeds TTS-1 HD in naturalness but costs 4× more.
Use cases
- Final narration and voice-over content where quality is paramount
- Podcast production and audio content where TTS-1 quality isn't sufficient
- Accessible content with premium audio requirements
Limitations
- Only 6 preset voices — same constraint as TTS-1
- No voice cloning or SSML support
- 2× cost of TTS-1; may be overkill for many use cases