← All models

OpenAI Whisper-1

Audio Deprecated Full comparison ↗

openai/whisper-1 · by OpenAI · encoder-decoder-transformer

Pricing — 1 offering(s)

Audio minutes

  • $0.006 / minute Current 2023-03-01 → 2027-02-26
    SCO-608 re-verification: unchanged at $0.006/minute. openai.com/api/pricing/ 403s to automated fetches…

    SCO-608 re-verification: unchanged at $0.006/minute. openai.com/api/pricing/ 403s to automated fetches (Cloudflare bot check, not dead); confirmed instead via developers.openai.com/api/docs/pricing, which lists the same figure. SCO-617: citation URL itself updated to match — the old URL now redirects to an unrelated ChatGPT Business seat-pricing page. SCO-672: effective_to is OpenAI's announced shutdown date (deprecations page, notice of 2026-08-26); the price stays current until then. It has left the docs pricing page but the model page still lists $0.006/minute.

Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.

See how OpenAI Whisper-1 fits into a cost-aware routing setup

See how →

Capability profile

How OpenAI Whisper-1 rates across core capability dimensions, with the task-level evidence behind each rating.

transcription accuracy Moderate

Strong on clean audio; degrades with background noise, heavy accents, or overlapping speech. OpenAI's newer GPT-4o-based transcription models surpass it on difficult audio. Best used where audio quality is controlled. Supplementary independent data point added 2026-08-19 (SCO-462 sweep): the Artificial Analysis Speech-to-Text leaderboard (artificialanalysis.ai/speech-to-text, AA-WER methodology, a different real-world dataset than the LibriSpeech figures already cited below) lists "Whisper Large v2" via OpenAI at 4.1% WER — solidly mid-to-back of the field on that benchmark (leaders like ElevenLabs Scribe v2 and Microsoft's MAI-Transcribe sit at 2.2-2.6%), consistent with this doc's existing "moderate" rating.

language support Strong

99 languages with automatic language detection. Includes a translation endpoint that transcribes audio directly to English — unique among the major hosted STT APIs.

speaker diarisation Weak

No native diarization. Word-level timestamps are available but speaker segmentation requires external post-processing (e.g. pyannote.audio).

real time streaming Weak

File-upload API only via /v1/audio/transcriptions and /v1/audio/translations. Latency scales with file duration; not suitable for real-time applications.

custom vocabulary Weak

No native custom vocabulary or keyword boosting. Prompt injection (prepending context text) provides limited guidance but is not reliable.

Benchmarks

Benchmark Score Config Source
WER (LibriSpeech clean) 2.7 % Whisper large-v2 paper result. API-hosted Whisper-1 uses this model checkpoint. source ↗
WER (LibriSpeech other) 5.2 % Whisper large-v2 on the noisier LibriSpeech Other split. source ↗

Operator guidance

Choose Whisper-1 when broad language coverage (99 languages) or direct audio-to-English translation matters. At $0.006/min it is cost-effective for batch workloads, but Nova-3 and Universal-2 both offer better accuracy on difficult audio. Not suitable for real-time streaming; for that use Deepgram.

Use cases

Limitations

Citations