OpenAI Whisper-1
Audio Deprecated Full comparison ↗ openai/whisper-1 · by OpenAI
· encoder-decoder-transformer
Pricing — 1 offering(s)
Audio minutes
- $0.006 / minute
Current
2023-03-01 → 2027-02-26
SCO-608 re-verification: unchanged at $0.006/minute. openai.com/api/pricing/ 403s to automated fetches…
SCO-608 re-verification: unchanged at $0.006/minute. openai.com/api/pricing/ 403s to automated fetches (Cloudflare bot check, not dead); confirmed instead via developers.openai.com/api/docs/pricing, which lists the same figure. SCO-617: citation URL itself updated to match — the old URL now redirects to an unrelated ChatGPT Business seat-pricing page. SCO-672: effective_to is OpenAI's announced shutdown date (deprecations page, notice of 2026-08-26); the price stays current until then. It has left the docs pricing page but the model page still lists $0.006/minute.
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how OpenAI Whisper-1 fits into a cost-aware routing setup
See how →Capability profile
How OpenAI Whisper-1 rates across core capability dimensions, with the task-level evidence behind each rating.
Strong on clean audio; degrades with background noise, heavy accents, or overlapping speech. OpenAI's newer GPT-4o-based transcription models surpass it on difficult audio. Best used where audio quality is controlled. Supplementary independent data point added 2026-08-19 (SCO-462 sweep): the Artificial Analysis Speech-to-Text leaderboard (artificialanalysis.ai/speech-to-text, AA-WER methodology, a different real-world dataset than the LibriSpeech figures already cited below) lists "Whisper Large v2" via OpenAI at 4.1% WER — solidly mid-to-back of the field on that benchmark (leaders like ElevenLabs Scribe v2 and Microsoft's MAI-Transcribe sit at 2.2-2.6%), consistent with this doc's existing "moderate" rating.
99 languages with automatic language detection. Includes a translation endpoint that transcribes audio directly to English — unique among the major hosted STT APIs.
No native diarization. Word-level timestamps are available but speaker segmentation requires external post-processing (e.g. pyannote.audio).
File-upload API only via /v1/audio/transcriptions and /v1/audio/translations. Latency scales with file duration; not suitable for real-time applications.
No native custom vocabulary or keyword boosting. Prompt injection (prepending context text) provides limited guidance but is not reliable.
Benchmarks
Operator guidance
Choose Whisper-1 when broad language coverage (99 languages) or direct audio-to-English translation matters. At $0.006/min it is cost-effective for batch workloads, but Nova-3 and Universal-2 both offer better accuracy on difficult audio. Not suitable for real-time streaming; for that use Deepgram.
Use cases
- Batch transcription of multilingual audio across 99 languages
- Transcription and translation in a single call (audio → English text)
- Research and experimentation where OpenAI platform consolidation matters
- Clean audio transcription where noise robustness is not a requirement
Limitations
- No streaming endpoint — batch only, file upload up to 25MB
- Degrades on noisy, accented, or overlapping audio
- No native speaker diarization
- No custom vocabulary support