ElevenLabs Eleven v3
Audio Available Full comparison ↗ elevenlabs/eleven-v3 · by ElevenLabs
· neural-tts
Pricing — 1 offering(s)
Text characters
- $100.00 / 1M characters
Historical
2026-07-30 → 2026-09-29
ElevenLabs' dedicated API pricing page lists "Multilingual v2 / v3" as one combined row at $0.10 per 1K…
ElevenLabs' dedicated API pricing page lists "Multilingual v2 / v3" as one combined row at $0.10 per 1K characters ($100/1M) — v3 is priced identically to Multilingual v2 on the pay-as-you-go API, not separately broken out. effective_from is the date this entry was sourced, not a confirmed price-change date.
- $80.00 / 1M characters
Current
2026-09-30 → present
SCO-672: elevenlabs.io/pricing/api now lists "v3" at $0.08 per 1K characters ($80/1M). That's the list price…
SCO-672: elevenlabs.io/pricing/api now lists "v3" at $0.08 per 1K characters ($80/1M). That's the list price: this card has no promo tag and no struck price (only v4 and v4 Turbo carry the "72% off until Oct 12" promo). SCO-346 read $0.10 per 1K from the same page on 2026-07-30, so ElevenLabs cut the rate between those dates. The exact change date is unknown (no archived snapshot was reachable), so effective_from is the date this was verified.
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how ElevenLabs Eleven v3 fits into a cost-aware routing setup
See how →Capability profile
How ElevenLabs Eleven v3 rates across core capability dimensions, with the task-level evidence behind each rating.
ElevenLabs' most expressive model — Audio Tags ([excited], [whispers], [sighs], etc.) give fine-grained emotional/delivery control beyond Multilingual v2's baseline. ElevenLabs' own GA testing reports a 68% error-rate reduction (15.3% to 4.9%) on numbers, symbols, and specialised notation across 8 languages vs the alpha release.
70+ languages per ElevenLabs' own v3 page — more than double Multilingual v2's 29.
Same 5,000+ voice library as Multilingual v2, plus a dialogue mode that weaves multiple voices into one multi-speaker generation, matching prosody and emotional range across speakers.
No latency figures published specific to v3 — not vendor-confirmed. Inferred comparable to Multilingual v2 (~400ms first-byte) as the shared base model; Audio Tags and dialogue-mode processing add overhead, so v3 is unlikely to be faster.
Full ElevenLabs voice cloning, same as Multilingual v2 and Flash v2.5.
Benchmarks
| Benchmark | Score | Config | Source |
|---|---|---|---|
| Pronunciation/notation accuracy improvement vs Eleven v3 alpha (vendor-reported) | 68 % | Error rate reduced from 15.3% to 4.9% across 27 categories (currency, chemical formulas, sports scores, phone numbers, etc.) in 8 languages, per ElevenLabs' own GA announcement. Not independently verified. | source ↗ |
Operator guidance
Choose Eleven v3 over Multilingual v2 when a script needs explicit emotional/delivery direction (Audio Tags) or multi-speaker dialogue — neither of which Multilingual v2 supports, at the same $80/1M-character price (SCO-672). Eleven v4 (2026-09-28) now supersedes v3 at the same list price (SCO-670). Route latency-sensitive, real-time applications to Flash v2.5 instead; v3's audio-tag and dialogue features are not oriented toward low-latency streaming.
Use cases
- Expressive, emotionally-directed narration where Audio Tags control delivery
- Multi-speaker dialogue and conversational audio via dialogue mode
- Premium content where directed emotional performance matters more than low latency
Limitations
- No independently published latency figures — not confirmed suitable for real-time/streaming use
- An independent overall-quality signal is now available (2026-08-19, SCO-462 sweep): the Artificial Analysis Text to Speech Arena (artificialanalysis.ai/text-to-speech/leaderboard/provider-voice, a human-preference Elo leaderboard) ranks Eleven v3 #12 of 98 voices tracked, Elo 1179. This is a general quality signal, not latency-specific, so the latency gap above is unaffected
- Credit-based pricing; effective rate is plan-dependent
- Audio Tag behavior is 'somewhat voice and context dependent' per ElevenLabs' own documentation — not guaranteed uniform across all voices