← All models

AssemblyAI Universal-2

Audio Available Full comparison ↗

assemblyai/universal-2 · by AssemblyAI · end-to-end-asr

Pricing — 1 offering(s)

Audio minutes

  • $0.0025 / minute Current 2024-06-01 → present
    SCO-608 re-verification: unchanged, $0.15/hr ($0.0025/min) for `universal-2` still current on the live…

    SCO-608 re-verification: unchanged, $0.15/hr ($0.0025/min) for `universal-2` still current on the live pricing page.

Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.

See how AssemblyAI Universal-2 fits into a cost-aware routing setup

See how →

Capability profile

How AssemblyAI Universal-2 rates across core capability dimensions, with the task-level evidence behind each rating.

transcription accuracy Strong

Top-tier accuracy on diverse audio including accents, low-quality recordings, and domain-specific terminology. Outperforms Whisper on difficult audio conditions; competitive with Deepgram Nova-3 on general audio.

language support Moderate

English-optimised with growing multilingual support. Best accuracy on English content; check AssemblyAI docs for the current supported language list.

speaker diarisation Strong

Built-in diarization with automatic speaker count detection. Supports speaker labelling and auto-chapters. Available as a feature flag; add-on pricing may apply at higher usage tiers.

real time streaming Moderate

Real-time streaming via WebSocket at ~500ms latency. Competent for most streaming use cases but higher latency than Deepgram's streaming endpoint.

custom vocabulary Strong

Word boost feature supports custom vocabulary lists to improve accuracy on domain-specific terms, names, and acronyms.

Benchmarks

Benchmark Score Config Source
WER (LibriSpeech) — AssemblyAI does not publish LibriSpeech WER benchmarks. Positioned as competitive with Nova-3 on general audio; independent comparisons show strong accuracy on diverse accents. —
Real-time streaming latency (vendor-reported) 500 ms Approximate WebSocket latency. Higher than Deepgram Nova-3's ~300ms. source ↗

Operator guidance

The best-value option for batch STT. At $0.0025/min ($0.15/hr) it is the lowest rate among major providers. Rich built-in features (diarization, chapters, PII redaction) reduce the need for post-processing. For streaming with the lowest latency, prefer Deepgram Nova-3. Note: pricing increases 10% from 2026-07-01 for in-region requests unless model_region=global is set.

Use cases

Limitations

Citations