Compare Google Cloud STT Chirp 2
Google Cloud STT Chirp 2’s pricing, architecture, and capability ratings side by side with up to three other AI audio models. Pick models below — your selection is saved in the URL and is shareable. Google Cloud STT Chirp 2 full profile ↗
Turn this comparison into a cost-aware routing setup
See how →What can I compare on this page?
Google Cloud STT Chirp 2's pricing, architecture, and capability ratings side by side with up to three other AI audio models. Add or remove comparison models with the selector below — the selection is saved in the page URL, so a specific comparison is shareable.
Is Google Cloud STT Chirp 2 or Azure Speech-to-Text cheaper?
Google Cloud STT Chirp 2 is cheaper: Google Cloud STT Chirp 2 is from $0.016 / minute, Azure Speech-to-Text is $0.0167 / minute.
How does Google Cloud STT Chirp 2 compare to Azure Speech-to-Text?
Modelglass rates both models across 5 capability dimensions. Azure Speech-to-Text rates higher on Custom vocabulary.
Compare with (up to 3)
| Google Cloud STT Chirp 2 base | Amazon Transcribe | AssemblyAI Universal-2 | Azure Speech-to-Text | Deepgram Nova-3 | MAI-Transcribe | OpenAI Whisper-1 (deprecated) | Whisper Large v3 | |
|---|---|---|---|---|---|---|---|---|
| Price | from $0.016 / minute | from $0.006 / minute | $0.0025 / minute | $0.0167 / minute | from $0.0043 / minute | $0.006 / minute | $0.006 / minute | $0.00185 / minute |
| Creator | Amazon | AssemblyAI | Microsoft | Deepgram | Microsoft AI | OpenAI | OpenAI | |
| Architecture | end-to-end-asr | end-to-end-asr | end-to-end-asr | end-to-end-asr | end-to-end-asr | llm-based-asr | encoder-decoder-transformer | encoder-decoder-transformer |
| Released | 2024-03 | 2018-04 | 2024-06 | 2017-06 | 2025-01 | 2026-04 | 2022-09 | 2023-11 |
| Generation | — | — | — | — | — | — | Previous | Current |
| Strong | Strong | Strong | Strong | Strong | Strong | Moderate | Strong | |
| Word error rate on diverse real-world audio — covering accented speech, background noise, domain-specific vocabulary, and low-quality recordings. What each rating means here
| ||||||||
| Strong | Strong | Moderate | Strong | Moderate | Moderate | Strong | Strong | |
| Number and quality of supported input languages. Strong coverage means reliable transcription across many major and minor languages. What each rating means here
| ||||||||
| Strong | Strong | Strong | Strong | Strong | Weak | Weak | Weak | |
| Ability to identify and label different speakers in a recording — critical for meeting transcription, interviews, and multi-participant audio. What each rating means here
| ||||||||
| Strong | Strong | Moderate | Strong | Strong | Weak | Weak | Moderate | |
| Latency and accuracy when transcribing live audio streams — lower latency and stable partial transcripts are key for interactive applications. What each rating means here
| ||||||||
| Moderate | Strong | Strong | Strong | Strong | Moderate | Weak | Weak | |
| Support for domain-specific terms, product names, and acronyms via word-boost, custom dictionaries, or fine-tuning. What each rating means here
| ||||||||
Capability ratings are an expert synthesis across benchmarks, community evaluations, and provider documentation. “—” means no profile data for that dimension.