← All models

GPT-4o Audio Preview

Audio Preview Full comparison ↗

openai/gpt-4o-audio-preview · by OpenAI · audio-native-transformer

Pricing — 1 offering(s)

Audio input tokens

  • $40.00 / 1M tokens (input) Historical 2024-10-01 → present
  • $80.00 / 1M tokens (output) Current 2024-10-01 → present

Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.

See how GPT-4o Audio Preview fits into a cost-aware routing setup

See how →

Capability profile

voice naturalness strong
language support strong
voice variety moderate
streaming latency strong
cloning support weak

Benchmarks

Benchmark Score Config Source
MOS (naturalness) No independently published MOS score found for gpt-4o-audio-preview specifically as of verification (2026-07-13). Re-checked 2026-08-19 (SCO-462 sweep) against the Artificial Analysis Text to Speech Arena (artificialanalysis.ai/text-to-speech/leaderboard/provider-voice, 98 voices tracked) — not found (only a distinct "GPT-Realtime-2" OpenAI entry appears, not this model). Confirmed still genuinely unavailable.

Operator guidance

OpenAI positions gpt-audio-1.5 as the GA successor to this model at a lower audio-token price ($32/$64 vs $40/$80 per 1M) — prefer gpt-audio-1.5 for new integrations. Choose this entry specifically only if you need GPT-4o's particular underlying model behaviour rather than the newer gpt-audio-1.5 model. For pure text-to-speech with no audio-input need, tts-1-hd is simpler and character-billed rather than token-billed.

Use cases

Limitations

Citations