GPT-4o Audio Preview
Audio PreviewCompare models
ComparingGPT-4o Audio Previewwith
Pick a model above to see the comparison.
openai/gpt-4o-audio-preview · by OpenAI · audio-native-transformer
Pricing — 1 offering(s)
Audio input tokens
- $40.00 / 1M tokens (input) Historical 2024-10-01 → present
- $80.00 / 1M tokens (output) Current 2024-10-01 → present
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
Capability profile
voice naturalness strong
language support strong
voice variety moderate
streaming latency strong
cloning support weak
Benchmarks
| Benchmark | Score | Config | Source |
|---|---|---|---|
| MOS (naturalness) | — | No independently published MOS score found for gpt-4o-audio-preview specifically as of verification (2026-07-13). | — |
Operator guidance
OpenAI positions gpt-audio-1.5 as the GA successor to this model at a lower audio-token price ($32/$64 vs $40/$80 per 1M) — prefer gpt-audio-1.5 for new integrations. Choose this entry specifically only if you need GPT-4o's particular underlying model behaviour rather than the newer gpt-audio-1.5 model. For pure text-to-speech with no audio-input need, tts-1-hd is simpler and character-billed rather than token-billed.
Use cases
- Conversational voice applications needing both audio understanding and audio response in one model call
- Use cases where GPT-4o's reasoning/instruction-following should directly drive speech output, not a separate TTS pass over generated text
- Migrating off this in favour of gpt-audio-1.5 unless a specific reason to stay on GPT-4o's underlying model requires it
Limitations
- Still labelled preview by OpenAI, not GA
- Token-billed pricing is harder to estimate cost for than tts-1's flat per-character rate — total cost depends on how many audio tokens a given utterance encodes to, which OpenAI does not publish a fixed ratio for
- No voice cloning or custom voice creation
- Superseded by gpt-audio-1.5 at a lower price for equivalent capability