← All models

Cartesia Sonic

Audio Available Full comparison ↗

cartesia/sonic · by Cartesia · state-space-model

Pricing — 1 offering(s)

Text characters

  • $65.00 / 1M characters Historical 2024-04-01 → 2026-09-19

    Verify exact rate at pricing page; Cartesia has updated pricing since launch.

  • $50.00 / 1M characters Current 2026-09-20 → present
    SCO-608 re-verification: real price change (down from $65 to $50/1M chars) plus a model-versioning update…

    SCO-608 re-verification: real price change (down from $65 to $50/1M chars) plus a model-versioning update — Cartesia's pricing page now names the model "Sonic-3.6", not plain "Sonic" (this entry's model.id/slug are unchanged for continuity; worth a generation-classifier check on whether a dedicated Sonic-3.6 entry is warranted). Cartesia has fully moved to a credits/subscription model with no PAYG per-character rate on the public pricing page itself — confirmed the conversion via Cartesia's own docs (docs.cartesia.ai/pricing): standard TTS is 1 credit per character. Rate here uses the Pro plan ($5/mo, 100K credits → $0.00005/credit = $50/1M chars), matching this registry's existing convention of pricing multi-tier consumer platforms off their standard/individual developer tier (see v4-suno). Other tiers differ: Startup ($49/mo, 1.25M credits) ≈ $39.20/1M chars; Scale ($299/mo, 8M credits) ≈ $37.38/1M chars — cheaper at higher committed volume, not recorded as separate tiers here.

Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.

See how Cartesia Sonic fits into a cost-aware routing setup

See how →

Capability profile

How Cartesia Sonic rates across core capability dimensions, with the task-level evidence behind each rating.

voice naturalness Moderate

Natural for a latency-optimised model; not best-in-class on prosody richness vs ElevenLabs Multilingual v2, but strong for real-time conversational AI.

language support Moderate

English-first; multilingual support is expanding but not yet at the breadth of Azure (140+) or Google (40+).

voice variety Moderate

Library of preset voices plus voice cloning (Sonic English). Smaller voice library than ElevenLabs or PlayHT.

streaming latency Strong

Sub-50ms first-audio-byte via WebSocket streaming — fastest among major TTS providers. The SSM architecture enables constant-time per-token generation.

cloning support Moderate

Voice cloning from reference audio supported. Quality is solid but not at the level of ElevenLabs Professional Clone.

Benchmarks

Benchmark Score Config Source
First-audio-byte latency (vendor-reported) 50 ms Vendor-reported sub-50ms target. Not independently verified against ElevenLabs Flash v2.5 (~75ms) under identical conditions. source ↗

Operator guidance

The primary choice when streaming latency is the dominant constraint. At ~$65/1M chars it costs more than ElevenLabs Flash v2.5 (~$50/1M) for only a modest quality difference, but the sub-50ms latency floor is architecturally superior for real-time voice agents. For batch or low-latency-tolerant TTS, ElevenLabs or PlayHT offer better value. Verify current pricing at cartesia.ai/pricing.

Use cases

Limitations

Citations