AssemblyAI Universal-2
Audio Available Full comparison ↗ assemblyai/universal-2 · by AssemblyAI
· end-to-end-asr
Pricing — 1 offering(s)
Audio minutes
- $0.0025 / minute
Current
2024-06-01 → present
SCO-608 re-verification: unchanged, $0.15/hr ($0.0025/min) for `universal-2` still current on the live…
SCO-608 re-verification: unchanged, $0.15/hr ($0.0025/min) for `universal-2` still current on the live pricing page.
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how AssemblyAI Universal-2 fits into a cost-aware routing setup
See how →Capability profile
How AssemblyAI Universal-2 rates across core capability dimensions, with the task-level evidence behind each rating.
Top-tier accuracy on diverse audio including accents, low-quality recordings, and domain-specific terminology. Outperforms Whisper on difficult audio conditions; competitive with Deepgram Nova-3 on general audio.
English-optimised with growing multilingual support. Best accuracy on English content; check AssemblyAI docs for the current supported language list.
Built-in diarization with automatic speaker count detection. Supports speaker labelling and auto-chapters. Available as a feature flag; add-on pricing may apply at higher usage tiers.
Real-time streaming via WebSocket at ~500ms latency. Competent for most streaming use cases but higher latency than Deepgram's streaming endpoint.
Word boost feature supports custom vocabulary lists to improve accuracy on domain-specific terms, names, and acronyms.
Benchmarks
| Benchmark | Score | Config | Source |
|---|---|---|---|
| WER (LibriSpeech) | — | AssemblyAI does not publish LibriSpeech WER benchmarks. Positioned as competitive with Nova-3 on general audio; independent comparisons show strong accuracy on diverse accents. | — |
| Real-time streaming latency (vendor-reported) | 500 ms | Approximate WebSocket latency. Higher than Deepgram Nova-3's ~300ms. | source ↗ |
Operator guidance
The best-value option for batch STT. At $0.0025/min ($0.15/hr) it is the lowest rate among major providers. Rich built-in features (diarization, chapters, PII redaction) reduce the need for post-processing. For streaming with the lowest latency, prefer Deepgram Nova-3. Note: pricing increases 10% from 2026-07-01 for in-region requests unless model_region=global is set.
Use cases
- Batch transcription at the lowest per-minute rate among major APIs
- Meeting transcription with speaker diarization and auto-chapters
- Content moderation and sentiment analysis pipelines
- Podcast and interview transcription with rich post-processing
Limitations
- English-first; multilingual coverage is narrower than Whisper
- Streaming latency (~500ms) is higher than Deepgram's offering
- Some advanced features (PII redaction, sentiment) add to the per-minute cost