Inworld TTS-2
Audio Available Full comparison ↗ inworld/tts-2 · by Inworld AI
· neural-tts
Pricing — 1 offering(s)
Text characters (On-Demand, no commitment)
- $25.00 / 1M characters
Current
2026-09-30 → present
inworld.ai/pricing, On-Demand plan (no commitment): "TTS-2 $25/1M chars". Monthly plans discount it (Creator…
inworld.ai/pricing, On-Demand plan (no commitment): "TTS-2 $25/1M chars". Monthly plans discount it (Creator $20/1M, Builder $17.50/1M, ...); the On-Demand rate is recorded as the pay-as-you-go comparison point, the same convention as tts-1-5-max-inworld. effective_from is the date this entry was sourced, not a launch or price-change date (SCO-672).
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Inworld TTS-2 fits into a cost-aware routing setup
See how →Capability profile
How Inworld TTS-2 rates across core capability dimensions, with the task-level evidence behind each rating.
200+ languages and locales in two tiers (15 most extensively evaluated, the rest fully supported), per Inworld's docs; vendor-reported.
No independent quality ranking found yet for TTS-2 (its predecessor TTS-1.5 Max ranked #8 of 98 on the Artificial Analysis TTS Arena). Not rated until one exists.
Instant voice cloning, per Inworld's docs; cloning quality not independently assessed.
100ms P90 time to first audio byte, measured server-side, per Inworld's docs; vendor-reported, excludes network latency.
Ratings are estimated — limited independent data is available for this model.
Benchmarks
| Benchmark | Score | Config | Source |
|---|---|---|---|
| P90 time to first audio byte (vendor-reported) | 100 ms | Server-side measurement, excludes network latency. Not independently verified. | source ↗ |
Operator guidance
Choose Inworld TTS-2 for low-cost ($25/1M on-demand), steerable realtime TTS across many languages. Choose TTS-2 Flash when latency matters more than steering. Quality claims are vendor-reported until an independent ranking exists.
Use cases
- Production realtime voice agents that need steerable delivery
- Interactive media and companions
Limitations
- Latency and language figures are vendor-reported
- No SSML; steering uses Inworld's square-bracket instructions
- No independent quality ranking yet