← All models

ElevenLabs Multilingual v2

Audio Available Full comparison ↗

elevenlabs/multilingual-v2 · by ElevenLabs · neural-tts

Pricing — 1 offering(s)

Text characters

  • $120.00 / 1M characters Historical 2023-10-01 → 2026-07-29
  • $100.00 / 1M characters Historical 2026-07-30 → 2026-09-29
    Found stale while primary-sourcing pricing for a separate Eleven v3 registry addition (SCO-346): the previous…

    Found stale while primary-sourcing pricing for a separate Eleven v3 registry addition (SCO-346): the previous $120/1M entry was an effective rate derived from Scale-plan credit arithmetic, per its own note ("ElevenLabs does not publish a direct per-character API rate"). That's no longer true — elevenlabs.io/pricing/api now lists a direct pay-as-you-go rate, "Multilingual v2 / v3: $0.10 per 1K characters" ($100/1M), grouping v2 and v3 at one shared rate. effective_to on the prior entry is the date before this one was first verified, not a confirmed price-change date.

  • $80.00 / 1M characters Current 2026-09-30 → present
    SCO-672: elevenlabs.io/pricing/api now lists "v2 Multilingual" at $0.08 per 1K characters ($80/1M). That's…

    SCO-672: elevenlabs.io/pricing/api now lists "v2 Multilingual" at $0.08 per 1K characters ($80/1M). That's the list price: this card has no promo tag and no struck price (only v4 and v4 Turbo carry the "72% off until Oct 12" promo). SCO-346 read $0.10 per 1K from the same page on 2026-07-30, so ElevenLabs cut the rate between those dates. The exact change date is unknown (no archived snapshot was reachable), so effective_from is the date this was verified.

Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.

See how ElevenLabs Multilingual v2 fits into a cost-aware routing setup

See how →

Capability profile

How ElevenLabs Multilingual v2 rates across core capability dimensions, with the task-level evidence behind each rating.

voice naturalness Strong

Best-in-class naturalness; often indistinguishable from human speech on English.

language support Strong

29 languages with high quality across European and East Asian languages.

voice variety Strong

Library of 5,000+ voices; professional voice cloning from as little as 1 minute of audio.

streaming latency Moderate

Higher latency than Flash v2.5 (~400ms first-byte). Not ideal for real-time.

cloning support Strong

Voice cloning, dubbing, speech-to-speech, emotion control via API.

Benchmarks

Benchmark Score Config Source
First-byte latency (vendor-reported) 400 ms Approximate; vendor-reported. Significantly higher than Flash v2.5's ~75ms. Not independently verified. source ↗
MOS (naturalness) — Consistently rated best-in-class in community comparisons. No formally published MOS score. —

Operator guidance

The reference choice for maximum voice quality. At ~$100/1M chars it is expensive but justified for content where audio quality drives user perception. Route latency-sensitive applications to Flash v2.5 instead. For simple utility TTS at scale, OpenAI TTS-1 is ~6.7× cheaper.

Use cases

Limitations

Citations