← All models

OpenAI TTS-1

Audio Available Full comparison ↗

openai/tts-1 · by OpenAI · neural-tts

Pricing — 1 offering(s)

Text characters

  • $15.00 / 1M characters Current 2023-11-06 → present
    SCO-608 re-verification: unchanged at $15/1M characters. openai.com/api/pricing/ 403s to automated fetches…

    SCO-608 re-verification: unchanged at $15/1M characters. openai.com/api/pricing/ 403s to automated fetches (Cloudflare bot check, not dead); confirmed instead via developers.openai.com/api/docs/pricing, which lists the same figure. SCO-617: citation URL itself updated to match — the old URL now redirects to an unrelated ChatGPT Business seat-pricing page.

Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.

See how OpenAI TTS-1 fits into a cost-aware routing setup

See how →

About OpenAI TTS-1

OpenAI TTS-1 is a proprietary neural text-to-speech model explicitly optimised for latency over fidelity. It offers 6 preset voices across 57 languages with chunked streaming through the API, and deliberately omits voice cloning, custom voice creation, and SSML — prosody control is limited to what the plain API exposes.

Its role in OpenAI's lineup is the low-cost option: at $15 per 1 million characters it is the cheapest major TTS API, with TTS-1 HD as the quality-tier sibling for voiceover and narration work where naturalness matters more than price or speed.

On independent measurement, the SCO-462 review (2026-08-19) found the Artificial Analysis Text-to-Speech Arena ranks "TTS-1" #34 of 98 voices tracked (Elo 1094) — mid-table, just behind TTS-1 HD (#30, Elo 1108) — which is consistent with Modelglass's "moderate" naturalness rating. There is no independently published MOS for it, and OpenAI does not disclose specific first-byte latency figures.

Capability profile

How OpenAI TTS-1 rates across core capability dimensions, with the task-level evidence behind each rating.

voice naturalness Moderate

Clear, natural-sounding speech for most content. Slight artifacts on complex phonemes. FILLED 2026-08-19 (SCO-462 sweep, was "no independently published MOS"): the Artificial Analysis Text to Speech Arena (artificialanalysis.ai/text-to-speech/leaderboard/provider-voice) ranks "TTS-1" #34 of 98 voices tracked, Elo 1094 — mid-table, just behind TTS-1 HD (#30, Elo 1108; see tts-1-hd.yaml), consistent with this doc's existing "moderate" rating.

language support Strong

Supports 57 languages. Quality varies by language; English is strongest.

voice variety Weak

Only 6 preset voices; no voice cloning or custom voice creation.

streaming latency Strong

Low latency; chunked streaming supported via the API.

cloning support Weak

No voice cloning, no fine-tuning, no SSML support. Preset voices only.

Benchmarks

Benchmark Score Config Source
MOS (naturalness) — No independently published MOS. Vendor positions as latency-optimised; community rates quality as moderate. —
First-byte latency — Vendor-described as low latency. Specific ms values not publicly disclosed by OpenAI. —

Operator guidance

Choose TTS-1 when cost and speed are the priority. At $15/1M chars it is the cheapest major TTS API. Upgrade to TTS-1 HD for quality-critical use (voice- overs, narration). Use ElevenLabs for natural voice quality or voice cloning.

Use cases

Limitations

Citations