Compare Whisper Large v3
Whisper Large v3’s pricing, architecture, and capability ratings side by side with up to three other AI audio models. Pick models below — your selection is saved in the URL and is shareable. Whisper Large v3 full profile ↗
Turn this comparison into a cost-aware routing setup
See how →What can I compare on this page?
Whisper Large v3's pricing, architecture, and capability ratings side by side with up to three other AI audio models. Add or remove comparison models with the selector below — the selection is saved in the page URL, so a specific comparison is shareable.
Is Whisper Large v3 or OpenAI Whisper-1 cheaper?
Whisper Large v3 is cheaper: Whisper Large v3 is $0.0019 / minute, OpenAI Whisper-1 is $0.0060 / minute.
How does Whisper Large v3 compare to OpenAI Whisper-1?
Modelglass rates both models across 5 capability dimensions. Whisper Large v3 rates higher on Transcription accuracy and Real-time streaming.
Compare with (up to 3)
| Whisper Large v3 base | Amazon Transcribe | AssemblyAI Universal-2 | Azure Speech-to-Text | Deepgram Nova-3 | Google Cloud STT Chirp 2 | MAI-Transcribe | OpenAI Whisper-1 | |
|---|---|---|---|---|---|---|---|---|
| Price | $0.0019 / minute | from $0.024 / minute | $0.0025 / minute | $0.017 / minute | from $0.0043 / minute | from $0.0060 / minute | $0.0060 / minute | $0.0060 / minute |
| Creator | OpenAI | Amazon | AssemblyAI | Microsoft | Deepgram | Microsoft AI | OpenAI | |
| Architecture | encoder-decoder-transformer | end-to-end-asr | end-to-end-asr | end-to-end-asr | end-to-end-asr | end-to-end-asr | llm-based-asr | encoder-decoder-transformer |
| Released | 2023-11 | 2018-04 | 2024-06 | 2017-06 | 2025-01 | 2024-03 | 2026-04 | 2022-09 |
| Generation | — | — | — | — | — | — | — | Previous |
| Strong | Strong | Strong | Strong | Strong | Strong | Strong | Moderate | |
| Word error rate on diverse real-world audio — covering accented speech, background noise, domain-specific vocabulary, and low-quality recordings. What each rating means here
| ||||||||
| Strong | Strong | Moderate | Strong | Moderate | Strong | Moderate | Strong | |
| Number and quality of supported input languages. Strong coverage means reliable transcription across many major and minor languages. What each rating means here
| ||||||||
| Weak | Strong | Strong | Strong | Strong | Strong | Weak | Weak | |
| Ability to identify and label different speakers in a recording — critical for meeting transcription, interviews, and multi-participant audio. What each rating means here
| ||||||||
| Moderate | Strong | Moderate | Strong | Strong | Strong | Weak | Weak | |
| Latency and accuracy when transcribing live audio streams — lower latency and stable partial transcripts are key for interactive applications. What each rating means here
| ||||||||
| Weak | Strong | Strong | Strong | Strong | Moderate | Moderate | Weak | |
| Support for domain-specific terms, product names, and acronyms via word-boost, custom dictionaries, or fine-tuning. What each rating means here
| ||||||||
Capability ratings are an expert synthesis across benchmarks, community evaluations, and provider documentation. “—” means no profile data for that dimension.