Compare MAI-Transcribe
MAI-Transcribe’s pricing, architecture, and capability ratings side by side with up to three other AI audio models. Pick models below — your selection is saved in the URL and is shareable. MAI-Transcribe full profile ↗
Turn this comparison into a cost-aware routing setup
See how →What can I compare on this page?
MAI-Transcribe's pricing, architecture, and capability ratings side by side with up to three other AI audio models. Add or remove comparison models with the selector below — the selection is saved in the page URL, so a specific comparison is shareable.
Is MAI-Transcribe or Azure Speech-to-Text cheaper?
MAI-Transcribe is cheaper: MAI-Transcribe is $0.0060 / minute, Azure Speech-to-Text is $0.017 / minute.
How does MAI-Transcribe compare to Azure Speech-to-Text?
Modelglass rates both models across 5 capability dimensions. Azure Speech-to-Text rates higher on Language support, Speaker diarisation, Real-time streaming, and Custom vocabulary.
Compare with (up to 3)
| MAI-Transcribe base | Amazon Transcribe | AssemblyAI Universal-2 | Azure Speech-to-Text | Deepgram Nova-3 | Google Cloud STT Chirp 2 | OpenAI Whisper-1 | Whisper Large v3 | |
|---|---|---|---|---|---|---|---|---|
| Price | $0.0060 / minute | from $0.024 / minute | $0.0025 / minute | $0.017 / minute | from $0.0043 / minute | from $0.0060 / minute | $0.0060 / minute | $0.0019 / minute |
| Creator | Microsoft AI | Amazon | AssemblyAI | Microsoft | Deepgram | OpenAI | OpenAI | |
| Architecture | llm-based-asr | end-to-end-asr | end-to-end-asr | end-to-end-asr | end-to-end-asr | end-to-end-asr | encoder-decoder-transformer | encoder-decoder-transformer |
| Released | 2026-04 | 2018-04 | 2024-06 | 2017-06 | 2025-01 | 2024-03 | 2022-09 | 2023-11 |
| Generation | — | — | — | — | — | — | Previous | — |
| Strong | Strong | Strong | Strong | Strong | Strong | Moderate | Strong | |
| Word error rate on diverse real-world audio — covering accented speech, background noise, domain-specific vocabulary, and low-quality recordings. What each rating means here
| ||||||||
| Moderate | Strong | Moderate | Strong | Moderate | Strong | Strong | Strong | |
| Number and quality of supported input languages. Strong coverage means reliable transcription across many major and minor languages. What each rating means here
| ||||||||
| Weak | Strong | Strong | Strong | Strong | Strong | Weak | Weak | |
| Ability to identify and label different speakers in a recording — critical for meeting transcription, interviews, and multi-participant audio. What each rating means here
| ||||||||
| Weak | Strong | Moderate | Strong | Strong | Strong | Weak | Moderate | |
| Latency and accuracy when transcribing live audio streams — lower latency and stable partial transcripts are key for interactive applications. What each rating means here
| ||||||||
| Moderate | Strong | Strong | Strong | Strong | Moderate | Weak | Weak | |
| Support for domain-specific terms, product names, and acronyms via word-boost, custom dictionaries, or fine-tuning. What each rating means here
| ||||||||
Capability ratings are an expert synthesis across benchmarks, community evaluations, and provider documentation. “—” means no profile data for that dimension.