Inworld TTS-2 Flash
Audio Available Full comparison ↗ inworld/tts-2-flash · by Inworld AI
· neural-tts
Pricing — 1 offering(s)
Text characters (On-Demand, no commitment)
- $15.00 / 1M characters
Current
2026-09-30 → present
inworld.ai/pricing, On-Demand plan (no commitment): "TTS-2 Flash $15/1M chars". Monthly plans discount it…
inworld.ai/pricing, On-Demand plan (no commitment): "TTS-2 Flash $15/1M chars". Monthly plans discount it (Creator $10/1M, Builder $9/1M, ...); the On-Demand rate is recorded as the pay-as-you-go comparison point, the same convention as tts-1-5-max-inworld. effective_from is the date this entry was sourced, not a launch or price-change date (SCO-672).
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Inworld TTS-2 Flash fits into a cost-aware routing setup
See how →Capability profile
How Inworld TTS-2 Flash rates across core capability dimensions, with the task-level evidence behind each rating.
200+ languages and locales in two tiers (15 most extensively evaluated, the rest fully supported), per Inworld's docs; vendor-reported.
No independent quality ranking found yet for TTS-2 (its predecessor TTS-1.5 Max ranked #8 of 98 on the Artificial Analysis TTS Arena). Not rated until one exists.
Instant voice cloning, per Inworld's docs; cloning quality not independently assessed.
20ms P90 time to first audio byte, measured server-side, "5x faster than inworld-tts-2", per Inworld's docs; vendor-reported.
Ratings are estimated — limited independent data is available for this model.
Benchmarks
| Benchmark | Score | Config | Source |
|---|---|---|---|
| P90 time to first audio byte (vendor-reported) | 20 ms | Server-side measurement, excludes network latency. Not independently verified. | source ↗ |
Operator guidance
Choose Inworld TTS-2 Flash for the lowest on-demand price in Inworld's line ($15/1M) and the lowest vendor-reported latency. Choose TTS-2 when steerable delivery matters.
Use cases
- High-volume, latency-critical voice agents
- Realtime conversational interfaces
Limitations
- Latency and language figures are vendor-reported
- No natural-language steering (TTS-2 has it)
- No independent quality ranking yet