Google Cloud STT Chirp 2
Audio Available Full comparison ↗ google-cloud/chirp-2 · by Google
· end-to-end-asr
· Built on google/universal-speech-model
Pricing — 1 offering(s)
Chirp 2 (enhanced accuracy)
- $0.012 / minute
Historical
2024-03-01 → 2026-09-19
Google STT v2 Chirp 2 model. First 60 minutes/month free.
- $0.016 / minute
Current
2026-09-20 → present
SCO-608 re-verification (browser-rendered — the page is heavily JS-templated and a plain fetch returns…
SCO-608 re-verification (browser-rendered — the page is heavily JS-templated and a plain fetch returns unusable minified JSON): Google restructured Speech-to-Text v2 pricing to a single volume- tiered "Standard recognition models" rate rather than separate per-model prices. Its own footnote states Standard¹ models "include: default, command_and_search, latest_short, latest_long, phone_call, video, chirp (Speech-to-Text V2 only)" — Chirp/Chirp 2 is now billed at the same Standard rate as every other v2 model, not a separate premium tier. $0.016/min is the first-500K-minutes- per-month band; volume discounts step down to $0.01 (500K-1M), $0.008 (1M-2M), $0.004/min (2M+) — the base/first-tier rate is recorded here, matching this registry's convention for other volume-tiered providers. Real, structural pricing-model change, not a data error — Chirp 2 is no longer more expensive than Standard.
Standard (previous-gen, V1 API, with data logging)
- $0.006 / minute
Historical
2016-03-01 → 2026-09-19
Google STT v1 standard models. First 60 minutes/month free.
- $0.016 / minute
Current
2026-09-20 → present
SCO-608 re-verification: same volume-tiered "Standard" rate as the chirp-2-audio-minutes tier above — see…
SCO-608 re-verification: same volume-tiered "Standard" rate as the chirp-2-audio-minutes tier above — see that tier's notes for the full V1/V2 restructuring detail. $0.016/min is the base (0-500K-minutes/month) band under the current v2 API; the legacy v1 API's "Standard" rate is now split into "with data logging" ($0.016/min) and "without data logging" ($0.024/min) instead of a single flat rate. SCO-673 (2026-09-30): this tier is the V1 API's with data logging price (sku 67F5-A183-E319, $0.016/min); the V1 rate without data logging is $0.024/min (sku 60AE-2FE3-C3D8) and deliberately has no tier of its own. The Chirp 2 tier above reads the V2 API's separate "Standard" SKU (3099-B70F-0949), so the two $0.016 figures come from different page cells.
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Google Cloud STT Chirp 2 fits into a cost-aware routing setup
See how →Capability profile
How Google Cloud STT Chirp 2 rates across core capability dimensions, with the task-level evidence behind each rating.
Strong accuracy across diverse accents, audio qualities, and domains. Chirp 2 improves on the original Chirp model especially on noisy and accented audio.
100+ languages with a single model. Quality holds well across major language families.
Native speaker diarisation supported via the v2 API. Multi-speaker detection with speaker labels.
Streaming transcription via gRPC bidirectional stream. Interim results and final transcripts supported.
Phrase hints and speech adaptation support for boosting domain-specific terms. Less flexible than Deepgram's Keywords API.
Benchmarks
| Benchmark | Score | Config | Source |
|---|---|---|---|
| WER (LibriSpeech) | — | Google does not publish LibriSpeech WER for Chirp 2 specifically. Internal benchmarks show improvements over Chirp 1 on diverse audio. | — |
Operator guidance
Best for teams in the Google Cloud ecosystem needing broad language coverage (100+ languages). At $0.012/min ($0.72/hr) it is cheaper than Azure ($1.00/hr) and Amazon Transcribe ($1.44/hr) but more expensive than Deepgram Nova-3 ($0.26/hr batch) and Universal-2 ($0.15/hr). The free tier (60 minutes/month) is useful for development. Choose Deepgram for the best streaming latency; Universal-2 for the lowest batch cost.
Use cases
- Multi-language transcription at scale via Google Cloud Platform
- Real-time transcription for 100+ language coverage
- Meeting and broadcast transcription with speaker diarisation
- Workflows already in the Google Cloud ecosystem
Limitations
- More expensive per minute than Deepgram and AssemblyAI for batch workloads
- Speech adaptation (custom vocabulary) is less flexible than Deepgram Keywords API
- gRPC streaming API has more setup complexity than REST