Deepgram Nova-3
Audio Available Full comparison ↗ deepgram/nova-3 · by Deepgram
· end-to-end-asr
Pricing — 1 offering(s)
Pre-recorded audio (batch)
- $0.0043 / minute
Current
2025-01-01 → present
SCO-608 re-verification: unchanged, PAYG monolingual rate still $0.0043/min ($0.0052/min for multilingual…
SCO-608 re-verification: unchanged, PAYG monolingual rate still $0.0043/min ($0.0052/min for multilingual, not separately tracked).
Streaming (real-time)
- $0.0077 / minute
Current
2025-01-01 → present
SCO-608 re-verification: unchanged at the PAYG regular rate of $0.0077/min. Deepgram is currently also…
SCO-608 re-verification: unchanged at the PAYG regular rate of $0.0077/min. Deepgram is currently also showing a limited-time promotional streaming rate ($0.0048/min monolingual) — not recorded here since it's explicitly time-limited, consistent with how this registry treats other promotional rates (see ling-3-0-flash-openrouter in modelglass-llm for precedent).
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Deepgram Nova-3 fits into a cost-aware routing setup
See how →About Deepgram Nova-3
Deepgram Nova-3 is a proprietary end-to-end ASR model trained on diverse audio — phone calls, meetings, broadcast media — rather than primarily on clean read speech. Over its predecessor Nova-2 it adds better punctuation, entity recognition and accent robustness. It is English-first; multilingual coverage is growing but its quality lags English noticeably.
Two things set it apart on the product side. Speaker diarization is built in via a `diarize=true` parameter and included in the base price, with no add-on surcharge unlike some competitors. And billing is per second with no minimum duration, which keeps cost predictable — though the streaming tier runs roughly 80% more expensive per minute than batch, and there is no native translation endpoint.
Deepgram positions Nova-3 as state-of-the-art on its own proprietary benchmarks (earnings calls, phone audio) and does not publish a LibriSpeech WER, so there is no standard independent accuracy number for it. The best-in-class streaming latency it cites — about 300ms round trip on Deepgram-hosted infrastructure — is vendor-reported and not independently verified.
Capability profile
How Deepgram Nova-3 rates across core capability dimensions, with the task-level evidence behind each rating.
State-of-the-art WER across accents and audio conditions. Particularly strong on phone-quality and noisy recordings. Improves on Nova-2 with better punctuation and entity recognition.
English-first with growing multilingual support. Coverage and accuracy are weaker than English; check Deepgram docs for the current language list.
Built-in diarization via diarize=true parameter. Accurate multi-speaker segmentation on meetings and conference audio. Included in base price — no add-on surcharge unlike some competitors.
Best-in-class streaming at ~300ms latency on Deepgram infrastructure. Per-second billing with no minimum duration. Supports interim results with is_final flags for display-while-speaking UX.
Keywords API supports custom word lists and phrase boosting for domain- specific terminology, product names, and proper nouns.
Benchmarks
| Benchmark | Score | Config | Source |
|---|---|---|---|
| Real-time streaming latency (vendor-reported) | 300 ms | Round-trip latency on Deepgram-hosted infrastructure. Vendor-reported, not independently verified. | source ↗ |
| WER (LibriSpeech) | — | Deepgram does not publish LibriSpeech WER. Nova-3 is positioned as state-of-the-art on proprietary benchmarks including earnings calls and phone audio. | — |
Operator guidance
The default choice for high-accuracy production STT. Best-in-class streaming latency makes it ideal for real-time apps; per-second billing (no rounding) keeps cost predictable. Use Whisper-1 for broad 99-language coverage or if you are already on the OpenAI platform. Use Universal-2 for the lowest per- minute rate on batch workloads.
Use cases
- Call centre and phone audio transcription at scale
- Real-time meeting transcription and captioning
- Voice assistants requiring low-latency incremental transcription
- Podcast and broadcast media transcription
Limitations
- Primarily English; multilingual quality lags significantly behind English
- Streaming tier is ~80% more expensive per minute than batch
- No native translation endpoint — transcription only