Cartesia Sonic
Audio Available Full comparison ↗ cartesia/sonic · by Cartesia
· state-space-model
Pricing — 1 offering(s)
Text characters
- $65.00 / 1M characters
Historical
2024-04-01 → 2026-09-19
Verify exact rate at pricing page; Cartesia has updated pricing since launch.
- $50.00 / 1M characters
Current
2026-09-20 → present
SCO-608 re-verification: real price change (down from $65 to $50/1M chars) plus a model-versioning update…
SCO-608 re-verification: real price change (down from $65 to $50/1M chars) plus a model-versioning update — Cartesia's pricing page now names the model "Sonic-3.6", not plain "Sonic" (this entry's model.id/slug are unchanged for continuity; worth a generation-classifier check on whether a dedicated Sonic-3.6 entry is warranted). Cartesia has fully moved to a credits/subscription model with no PAYG per-character rate on the public pricing page itself — confirmed the conversion via Cartesia's own docs (docs.cartesia.ai/pricing): standard TTS is 1 credit per character. Rate here uses the Pro plan ($5/mo, 100K credits → $0.00005/credit = $50/1M chars), matching this registry's existing convention of pricing multi-tier consumer platforms off their standard/individual developer tier (see v4-suno). Other tiers differ: Startup ($49/mo, 1.25M credits) ≈ $39.20/1M chars; Scale ($299/mo, 8M credits) ≈ $37.38/1M chars — cheaper at higher committed volume, not recorded as separate tiers here.
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Cartesia Sonic fits into a cost-aware routing setup
See how →Capability profile
How Cartesia Sonic rates across core capability dimensions, with the task-level evidence behind each rating.
Natural for a latency-optimised model; not best-in-class on prosody richness vs ElevenLabs Multilingual v2, but strong for real-time conversational AI.
English-first; multilingual support is expanding but not yet at the breadth of Azure (140+) or Google (40+).
Library of preset voices plus voice cloning (Sonic English). Smaller voice library than ElevenLabs or PlayHT.
Sub-50ms first-audio-byte via WebSocket streaming — fastest among major TTS providers. The SSM architecture enables constant-time per-token generation.
Voice cloning from reference audio supported. Quality is solid but not at the level of ElevenLabs Professional Clone.
Benchmarks
| Benchmark | Score | Config | Source |
|---|---|---|---|
| First-audio-byte latency (vendor-reported) | 50 ms | Vendor-reported sub-50ms target. Not independently verified against ElevenLabs Flash v2.5 (~75ms) under identical conditions. | source ↗ |
Operator guidance
The primary choice when streaming latency is the dominant constraint. At ~$65/1M chars it costs more than ElevenLabs Flash v2.5 (~$50/1M) for only a modest quality difference, but the sub-50ms latency floor is architecturally superior for real-time voice agents. For batch or low-latency-tolerant TTS, ElevenLabs or PlayHT offer better value. Verify current pricing at cartesia.ai/pricing.
Use cases
- Real-time conversational AI requiring sub-50ms voice response
- Interactive voice agents where latency is the primary constraint
- Low-latency TTS pipelines replacing transformers-based models
- Applications sensitive to streaming latency (voice assistants, phone bots)
Limitations
- English-first; limited multilingual coverage vs Azure or Google
- Smaller voice library than ElevenLabs or PlayHT
- Pricing has changed post-launch; verify current rate before use
- Early-stage company; pricing and feature stability less certain than established cloud providers