GPT-Audio 1.5
Audio AvailableCompare models
ComparingGPT-Audio 1.5with
Pick a model above to see the comparison.
openai/gpt-audio-1.5 · by OpenAI · audio-native-transformer
Pricing — 1 offering(s)
Audio input tokens
- $32.00 / 1M tokens (input) Historical 2026-07-13 → present
- $64.00 / 1M tokens (output) Current 2026-07-13 → present
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
Capability profile
voice naturalness strong
language support strong
voice variety moderate
streaming latency strong
cloning support weak
Benchmarks
| Benchmark | Score | Config | Source |
|---|---|---|---|
| MOS (naturalness) | — | No independently published MOS score found for gpt-audio-1.5 specifically as of verification (2026-07-13). | — |
Operator guidance
This is OpenAI's current recommended audio-native model — cheaper than gpt-4o-audio-preview at the same $2.50/$10 text-token rate, with lower $32/$64 (vs $40/$80) per-1M audio-token pricing. For pure text-to-speech with no audio-input need, tts-1-hd is simpler and character-billed rather than token-billed, which is easier to cost-estimate up front.
Use cases
- Default choice for new OpenAI audio-native integrations per OpenAI's own 'Default' labelling
- Conversational voice applications needing both audio understanding and audio response in one model call
- Lower-cost alternative to gpt-4o-audio-preview for the same bidirectional-audio capability
Limitations
- Token-billed pricing is harder to estimate cost for than tts-1's flat per-character rate — total cost depends on how many audio tokens a given utterance encodes to, which OpenAI does not publish a fixed ratio for
- No voice cloning or custom voice creation
- No published release date found on OpenAI's own model documentation as of verification (2026-07-13) — a third-party source claims April 2026, unconfirmed against a first-party OpenAI page