OpenAI TTS-1
Audio Available Full comparison ↗ openai/tts-1 · by OpenAI
· neural-tts
Pricing — 1 offering(s)
Text characters
- $15.00 / 1M characters
Current
2023-11-06 → present
SCO-608 re-verification: unchanged at $15/1M characters. openai.com/api/pricing/ 403s to automated fetches…
SCO-608 re-verification: unchanged at $15/1M characters. openai.com/api/pricing/ 403s to automated fetches (Cloudflare bot check, not dead); confirmed instead via developers.openai.com/api/docs/pricing, which lists the same figure. SCO-617: citation URL itself updated to match — the old URL now redirects to an unrelated ChatGPT Business seat-pricing page.
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how OpenAI TTS-1 fits into a cost-aware routing setup
See how →About OpenAI TTS-1
OpenAI TTS-1 is a proprietary neural text-to-speech model explicitly optimised for latency over fidelity. It offers 6 preset voices across 57 languages with chunked streaming through the API, and deliberately omits voice cloning, custom voice creation, and SSML — prosody control is limited to what the plain API exposes.
Its role in OpenAI's lineup is the low-cost option: at $15 per 1 million characters it is the cheapest major TTS API, with TTS-1 HD as the quality-tier sibling for voiceover and narration work where naturalness matters more than price or speed.
On independent measurement, the SCO-462 review (2026-08-19) found the Artificial Analysis Text-to-Speech Arena ranks "TTS-1" #34 of 98 voices tracked (Elo 1094) — mid-table, just behind TTS-1 HD (#30, Elo 1108) — which is consistent with Modelglass's "moderate" naturalness rating. There is no independently published MOS for it, and OpenAI does not disclose specific first-byte latency figures.
Capability profile
How OpenAI TTS-1 rates across core capability dimensions, with the task-level evidence behind each rating.
Clear, natural-sounding speech for most content. Slight artifacts on complex phonemes. FILLED 2026-08-19 (SCO-462 sweep, was "no independently published MOS"): the Artificial Analysis Text to Speech Arena (artificialanalysis.ai/text-to-speech/leaderboard/provider-voice) ranks "TTS-1" #34 of 98 voices tracked, Elo 1094 — mid-table, just behind TTS-1 HD (#30, Elo 1108; see tts-1-hd.yaml), consistent with this doc's existing "moderate" rating.
Supports 57 languages. Quality varies by language; English is strongest.
Only 6 preset voices; no voice cloning or custom voice creation.
Low latency; chunked streaming supported via the API.
No voice cloning, no fine-tuning, no SSML support. Preset voices only.
Benchmarks
| Benchmark | Score | Config | Source |
|---|---|---|---|
| MOS (naturalness) | — | No independently published MOS. Vendor positions as latency-optimised; community rates quality as moderate. | — |
| First-byte latency | — | Vendor-described as low latency. Specific ms values not publicly disclosed by OpenAI. | — |
Operator guidance
Choose TTS-1 when cost and speed are the priority. At $15/1M chars it is the cheapest major TTS API. Upgrade to TTS-1 HD for quality-critical use (voice- overs, narration). Use ElevenLabs for natural voice quality or voice cloning.
Use cases
- App TTS at scale where cost matters more than premium voice quality
- Assistants and chatbots needing fast speech output
- Multi-language content narration
Limitations
- Only 6 preset voices — no voice cloning or custom voice creation
- No SSML support; limited prosody control
- Audio quality noticeably below ElevenLabs Multilingual v2 on close listening