ElevenLabs Multilingual v2
Audio Available Full comparison ↗ elevenlabs/multilingual-v2 · by ElevenLabs
· neural-tts
Pricing — 1 offering(s)
Text characters
- $120.00 / 1M characters Historical 2023-10-01 → 2026-07-29
- $100.00 / 1M characters
Historical
2026-07-30 → 2026-09-29
Found stale while primary-sourcing pricing for a separate Eleven v3 registry addition (SCO-346): the previous…
Found stale while primary-sourcing pricing for a separate Eleven v3 registry addition (SCO-346): the previous $120/1M entry was an effective rate derived from Scale-plan credit arithmetic, per its own note ("ElevenLabs does not publish a direct per-character API rate"). That's no longer true — elevenlabs.io/pricing/api now lists a direct pay-as-you-go rate, "Multilingual v2 / v3: $0.10 per 1K characters" ($100/1M), grouping v2 and v3 at one shared rate. effective_to on the prior entry is the date before this one was first verified, not a confirmed price-change date.
- $80.00 / 1M characters
Current
2026-09-30 → present
SCO-672: elevenlabs.io/pricing/api now lists "v2 Multilingual" at $0.08 per 1K characters ($80/1M). That's…
SCO-672: elevenlabs.io/pricing/api now lists "v2 Multilingual" at $0.08 per 1K characters ($80/1M). That's the list price: this card has no promo tag and no struck price (only v4 and v4 Turbo carry the "72% off until Oct 12" promo). SCO-346 read $0.10 per 1K from the same page on 2026-07-30, so ElevenLabs cut the rate between those dates. The exact change date is unknown (no archived snapshot was reachable), so effective_from is the date this was verified.
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how ElevenLabs Multilingual v2 fits into a cost-aware routing setup
See how →Capability profile
How ElevenLabs Multilingual v2 rates across core capability dimensions, with the task-level evidence behind each rating.
Best-in-class naturalness; often indistinguishable from human speech on English.
29 languages with high quality across European and East Asian languages.
Library of 5,000+ voices; professional voice cloning from as little as 1 minute of audio.
Higher latency than Flash v2.5 (~400ms first-byte). Not ideal for real-time.
Voice cloning, dubbing, speech-to-speech, emotion control via API.
Benchmarks
| Benchmark | Score | Config | Source |
|---|---|---|---|
| First-byte latency (vendor-reported) | 400 ms | Approximate; vendor-reported. Significantly higher than Flash v2.5's ~75ms. Not independently verified. | source ↗ |
| MOS (naturalness) | — | Consistently rated best-in-class in community comparisons. No formally published MOS score. | — |
Operator guidance
The reference choice for maximum voice quality. At ~$100/1M chars it is expensive but justified for content where audio quality drives user perception. Route latency-sensitive applications to Flash v2.5 instead. For simple utility TTS at scale, OpenAI TTS-1 is ~6.7× cheaper.
Use cases
- Premium audiobooks, podcasts, and narration
- Voice cloning for consistent brand voice across content
- Dubbing and localisation workflows
- High-quality IVR systems where naturalness matters
Limitations
- Credit-based pricing; effective rate is plan-dependent
- Higher latency than Flash v2.5; not suitable for real-time streaming
- Directly published per-character API rate; may not reflect enterprise pricing