← All models

Google Cloud STT Chirp 2

Audio Available Full comparison ↗

google-cloud/chirp-2 · by Google · end-to-end-asr · Built on google/universal-speech-model

Pricing — 1 offering(s)

Chirp 2 (enhanced accuracy)

  • $0.012 / minute Historical 2024-03-01 → 2026-09-19

    Google STT v2 Chirp 2 model. First 60 minutes/month free.

  • $0.016 / minute Current 2026-09-20 → present
    SCO-608 re-verification (browser-rendered — the page is heavily JS-templated and a plain fetch returns…

    SCO-608 re-verification (browser-rendered — the page is heavily JS-templated and a plain fetch returns unusable minified JSON): Google restructured Speech-to-Text v2 pricing to a single volume- tiered "Standard recognition models" rate rather than separate per-model prices. Its own footnote states Standard¹ models "include: default, command_and_search, latest_short, latest_long, phone_call, video, chirp (Speech-to-Text V2 only)" — Chirp/Chirp 2 is now billed at the same Standard rate as every other v2 model, not a separate premium tier. $0.016/min is the first-500K-minutes- per-month band; volume discounts step down to $0.01 (500K-1M), $0.008 (1M-2M), $0.004/min (2M+) — the base/first-tier rate is recorded here, matching this registry's convention for other volume-tiered providers. Real, structural pricing-model change, not a data error — Chirp 2 is no longer more expensive than Standard.

Standard (previous-gen, V1 API, with data logging)

  • $0.006 / minute Historical 2016-03-01 → 2026-09-19

    Google STT v1 standard models. First 60 minutes/month free.

  • $0.016 / minute Current 2026-09-20 → present
    SCO-608 re-verification: same volume-tiered "Standard" rate as the chirp-2-audio-minutes tier above — see…

    SCO-608 re-verification: same volume-tiered "Standard" rate as the chirp-2-audio-minutes tier above — see that tier's notes for the full V1/V2 restructuring detail. $0.016/min is the base (0-500K-minutes/month) band under the current v2 API; the legacy v1 API's "Standard" rate is now split into "with data logging" ($0.016/min) and "without data logging" ($0.024/min) instead of a single flat rate. SCO-673 (2026-09-30): this tier is the V1 API's with data logging price (sku 67F5-A183-E319, $0.016/min); the V1 rate without data logging is $0.024/min (sku 60AE-2FE3-C3D8) and deliberately has no tier of its own. The Chirp 2 tier above reads the V2 API's separate "Standard" SKU (3099-B70F-0949), so the two $0.016 figures come from different page cells.

Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.

See how Google Cloud STT Chirp 2 fits into a cost-aware routing setup

See how →

Capability profile

How Google Cloud STT Chirp 2 rates across core capability dimensions, with the task-level evidence behind each rating.

transcription accuracy Strong

Strong accuracy across diverse accents, audio qualities, and domains. Chirp 2 improves on the original Chirp model especially on noisy and accented audio.

language support Strong

100+ languages with a single model. Quality holds well across major language families.

speaker diarisation Strong

Native speaker diarisation supported via the v2 API. Multi-speaker detection with speaker labels.

real time streaming Strong

Streaming transcription via gRPC bidirectional stream. Interim results and final transcripts supported.

custom vocabulary Moderate

Phrase hints and speech adaptation support for boosting domain-specific terms. Less flexible than Deepgram's Keywords API.

Benchmarks

Benchmark Score Config Source
WER (LibriSpeech) — Google does not publish LibriSpeech WER for Chirp 2 specifically. Internal benchmarks show improvements over Chirp 1 on diverse audio. —

Operator guidance

Best for teams in the Google Cloud ecosystem needing broad language coverage (100+ languages). At $0.012/min ($0.72/hr) it is cheaper than Azure ($1.00/hr) and Amazon Transcribe ($1.44/hr) but more expensive than Deepgram Nova-3 ($0.26/hr batch) and Universal-2 ($0.15/hr). The free tier (60 minutes/month) is useful for development. Choose Deepgram for the best streaming latency; Universal-2 for the lowest batch cost.

Use cases

Limitations

Citations