Skip to navigation
Inworld

Inworld

Models from Inworld

TTS 2 Flash
TTS 2 Flash
Latency-first realtime voice synthesis holding one voice identity across 200+ languages, tuned for high-volume streaming workloads.
TTS 2
TTS 2
Realtime voice model that takes plain-English delivery direction, holds one voice identity across 200+ languages, and streams with word timestamps.
TTS 1.5 Mini
TTS 1.5 Mini
Sub-130ms TTFB voice synthesis with 271+ voices across 15 languages, expressive prosody, and real-time SSE streaming for low-latency voice agents.
TTS 1.5 Max
TTS 1.5 Max
Broadcast-quality voice synthesis with rich expressive prosody, 271+ voices across 15 languages, and real-time SSE streaming with per-word timestamps.