Inworld TTS-1.5 Max
Audio Deprecated Full comparison ↗ inworld/tts-1-5-max · by Inworld AI
· neural-tts
Pricing — 1 offering(s)
Text characters (On-Demand, no commitment)
- $35.00 / 1M characters
Current
2026-07-30 → present
Inworld's Realtime TTS-1.5 Max is tiered by monthly spend commitment, not a single flat rate: On-Demand (no…
Inworld's Realtime TTS-1.5 Max is tiered by monthly spend commitment, not a single flat rate: On-Demand (no commitment) $35/1M chars, Creator $25/1M, Builder $22.50/1M, Developer $20/1M, Growth $17.50/1M, Enterprise custom (as low as $5/1M on their cheaper TTS-2 line at enterprise volume). The On-Demand rate is recorded here as the apples-to-apples pay-as-you-go comparison point, matching how other providers in this registry are priced. effective_from is the date this entry was sourced, not a confirmed price-change date.
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Inworld TTS-1.5 Max fits into a cost-aware routing setup
See how →Capability profile
How Inworld TTS-1.5 Max rates across core capability dimensions, with the task-level evidence behind each rating.
Positioned by Inworld as the higher-quality tier of the TTS-1.5 family (vs TTS-1.5 Mini), for realtime, production-grade voice agents. FILLED 2026-08-19 (SCO-462 sweep, was vendor-positioning only): independent corroboration now exists — the Artificial Analysis Text to Speech Arena (artificialanalysis.ai/text-to-speech/leaderboard/provider-voice, a human-preference Elo leaderboard) ranks "Inworld Realtime TTS 1.5 Max" #8 of 98 voices tracked, Elo 1196 — a genuinely strong independent result, ahead of ElevenLabs Eleven v3 (#12, Elo 1179) and Fish Audio S2.1 Pro (#18, Elo 1144).
15 languages, per Inworld's own product page — narrower than ElevenLabs (29-70+) or Fish Audio's S2 line (80+), though Inworld's positioning prioritises realtime latency over language breadth.
Inworld's product page doesn't publish an exact voice-library count, unlike ElevenLabs (5,000+) or Fish Audio.
P90 <250ms, median ~200ms first-chunk latency per Inworld's own published figures — among the fastest vendor-reported numbers in this registry; not independently verified against other providers under identical conditions.
Instant cloning from 5-15 seconds of reference audio; a professional cloning tier (30+ minutes of audio, via sales) is also offered.
Benchmarks
| Benchmark | Score | Config | Source |
|---|---|---|---|
| P90 first-chunk latency (vendor-reported) | 250 ms | Median ~200ms per the same source. Vendor-reported, not independently verified against other providers under identical conditions. | source ↗ |
Operator guidance
Choose Inworld TTS-1.5 Max when sub-250ms streaming latency for voice agents is the priority — its WebSocket-native architecture and published P90 latency figures are among the fastest vendor-reported numbers in this registry. Fish Audio S2 Pro/S2.1 Pro offer broader language coverage (80+) at a comparable price point if language breadth matters more than Inworld's realtime-first positioning.
Use cases
- Realtime, production-grade voice agents
- Audiobooks and interactive media
- Accessibility tools and language tutoring
- Content creation at scale
Limitations
- No SSML support mentioned in Inworld's documentation
- Narrower language support (15) than ElevenLabs or Fish Audio's S2 line
- No published voice-library count, unlike ElevenLabs or Fish Audio
- Latency figures specifically are still vendor-published, not independently verified — the Arena Elo cited on voice-naturalness is a general quality signal, not latency-specific