Compare GPT-4o Audio Preview
GPT-4o Audio Preview’s pricing, architecture, and capability ratings side by side with up to three other AI audio models. Pick models below — your selection is saved in the URL and is shareable. GPT-4o Audio Preview full profile ↗
Turn this comparison into a cost-aware routing setup
See how →What can I compare on this page?
GPT-4o Audio Preview's pricing, architecture, and capability ratings side by side with up to three other AI audio models. Add or remove comparison models with the selector below — the selection is saved in the page URL, so a specific comparison is shareable.
Is GPT-4o Audio Preview or GPT-Audio 1.5 cheaper?
GPT-Audio 1.5 is cheaper: GPT-4o Audio Preview is $0.096 / minute, GPT-Audio 1.5 is $0.077 / minute.
How does GPT-4o Audio Preview compare to GPT-Audio 1.5?
Modelglass rates both models across 5 capability dimensions. They rate evenly on the dimensions where both have profile data.
Compare with (up to 3)
| GPT-4o Audio Preview base | Amazon Polly | Azure Neural TTS | Cartesia Sonic | ElevenLabs Eleven v3 | ElevenLabs Flash v2.5 | ElevenLabs Multilingual v2 | Fish Audio S1 | Fish Audio S2 Pro | Fish Audio S2.1 Pro | Google Cloud TTS | GPT-Audio 1.5 | Grok Voice Think Fast 2.0 | Inworld TTS-1.5 Max | OpenAI TTS-1 | OpenAI TTS-1 HD | PlayHT 2.0 (retired) | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Price | $0.096 / minute | from $4.00 / 1M characters | from $4.00 / 1M characters | $65.00 / 1M characters | $100.00 / 1M characters | $50.00 / 1M characters | $100.00 / 1M characters | $15.00 / 1M characters | $15.00 / 1M characters | $15.00 / 1M characters | from $4.00 / 1M characters | $0.077 / minute | $0.080 / minute | $35.00 / 1M characters | $15.00 / 1M characters | $30.00 / 1M characters | — |
| Creator | OpenAI | Amazon | Microsoft | Cartesia | ElevenLabs | ElevenLabs | ElevenLabs | Fish Audio | Fish Audio | Fish Audio | OpenAI | xAI | Inworld AI | OpenAI | OpenAI | PlayHT | |
| Architecture | audio-native-transformer | neural-tts | neural-tts | state-space-model | neural-tts | neural-tts | neural-tts | neural-tts | neural-tts | neural-tts | neural-tts | audio-native-transformer | audio-native-transformer | neural-tts | neural-tts | neural-tts | neural-tts |
| Released | 2024-10 | 2016-11 | 2019-09 | 2024-04 | 2026-02 | 2024-06 | 2023-10 | 2025-11 | 2026 | 2026 | 2018-03 | — | 2026-07 | 2026-01 | 2023-11 | 2023-11 | 2023-07 |
| Generation | Previous | — | — | — | Current | — | Previous | Previous | Previous | Current | — | Current | Current | — | Previous | Previous | — |
| Strong | Moderate | Strong | Moderate | Strong | Moderate | Strong | Strong | Strong | Strong | Strong | Strong | Strong | Strong | Moderate | Strong | Strong | |
| How natural and human-like the synthesized speech sounds — covering prosody, pacing, intonation, and absence of robotic or mechanical artefacts. What each rating means here
| |||||||||||||||||
| Moderate | Moderate | Strong | Moderate | Strong | Strong | Strong | Strong | Strong | Strong | Strong | Moderate | Unknown | Unknown | Weak | Weak | Strong | |
| The range of distinct voices, accents, ages, and speaking styles available out of the box from the provider's voice library. What each rating means here
| |||||||||||||||||
| Weak | Weak | Moderate | Moderate | Strong | Strong | Strong | Strong | Strong | Strong | Weak | Weak | Unknown | Strong | Weak | Weak | Strong | |
| Ability to clone a custom voice from a short audio sample, enabling personalised or brand-consistent TTS output. What each rating means here
| |||||||||||||||||
| Strong | Strong | Moderate | Strong | Moderate | Strong | Moderate | Unknown | Strong | Strong | Moderate | Strong | Strong | Strong | Strong | Moderate | Moderate | |
| Time-to-first-audio-chunk when using the streaming endpoint — lower latency enables real-time conversational applications. What each rating means here
| |||||||||||||||||
| Strong | Strong | Strong | Moderate | Strong | Strong | Strong | Moderate | Strong | Strong | Strong | Strong | Strong | Moderate | Strong | Strong | Strong | |
| Number and quality of supported output languages. Strong coverage means high-quality synthesis across many major and minor languages. What each rating means here
| |||||||||||||||||
Capability ratings are an expert synthesis across benchmarks, community evaluations, and provider documentation. “—” means no profile data for that dimension.