Compare Deepgram Nova-3
Deepgram Nova-3’s pricing, architecture, and capability ratings side by side with up to three other AI audio models. Pick models below — your selection is saved in the URL and is shareable. Deepgram Nova-3 full profile ↗
Turn this comparison into a cost-aware routing setup
See how →What can I compare on this page?
Deepgram Nova-3's pricing, architecture, and capability ratings side by side with up to three other AI audio models. Add or remove comparison models with the selector below — the selection is saved in the page URL, so a specific comparison is shareable.
Is Deepgram Nova-3 or AssemblyAI Universal-2 cheaper?
AssemblyAI Universal-2 is cheaper: Deepgram Nova-3 is from $0.0043 / minute, AssemblyAI Universal-2 is $0.0025 / minute.
How does Deepgram Nova-3 compare to AssemblyAI Universal-2?
Modelglass rates both models across 5 capability dimensions. Deepgram Nova-3 rates higher on Real-time streaming.
Compare with (up to 3)
| Deepgram Nova-3 base | Amazon Transcribe | AssemblyAI Universal-2 | Azure Speech-to-Text | Google Cloud STT Chirp 2 | MAI-Transcribe | OpenAI Whisper-1 (deprecated) | Whisper Large v3 | |
|---|---|---|---|---|---|---|---|---|
| Price | from $0.0043 / minute | from $0.006 / minute | $0.0025 / minute | $0.0167 / minute | from $0.016 / minute | $0.006 / minute | $0.006 / minute | $0.00185 / minute |
| Creator | Deepgram | Amazon | AssemblyAI | Microsoft | Microsoft AI | OpenAI | OpenAI | |
| Architecture | end-to-end-asr | end-to-end-asr | end-to-end-asr | end-to-end-asr | end-to-end-asr | llm-based-asr | encoder-decoder-transformer | encoder-decoder-transformer |
| Released | 2025-01 | 2018-04 | 2024-06 | 2017-06 | 2024-03 | 2026-04 | 2022-09 | 2023-11 |
| Generation | — | — | — | — | — | — | Previous | Current |
| Strong | Strong | Strong | Strong | Strong | Strong | Moderate | Strong | |
| Word error rate on diverse real-world audio — covering accented speech, background noise, domain-specific vocabulary, and low-quality recordings. What each rating means here
| ||||||||
| Moderate | Strong | Moderate | Strong | Strong | Moderate | Strong | Strong | |
| Number and quality of supported input languages. Strong coverage means reliable transcription across many major and minor languages. What each rating means here
| ||||||||
| Strong | Strong | Strong | Strong | Strong | Weak | Weak | Weak | |
| Ability to identify and label different speakers in a recording — critical for meeting transcription, interviews, and multi-participant audio. What each rating means here
| ||||||||
| Strong | Strong | Moderate | Strong | Strong | Weak | Weak | Moderate | |
| Latency and accuracy when transcribing live audio streams — lower latency and stable partial transcripts are key for interactive applications. What each rating means here
| ||||||||
| Strong | Strong | Strong | Strong | Moderate | Moderate | Weak | Weak | |
| Support for domain-specific terms, product names, and acronyms via word-boost, custom dictionaries, or fine-tuning. What each rating means here
| ||||||||
Capability ratings are an expert synthesis across benchmarks, community evaluations, and provider documentation. “—” means no profile data for that dimension.