Inworld

Inworld
Models from Inworld

TTS 2 Flash
Latency-first realtime voice synthesis holding one voice identity across 200+ languages, tuned for high-volume streaming workloads.

TTS 2
Realtime voice model that takes plain-English delivery direction, holds one voice identity across 200+ languages, and streams with word timestamps.

TTS 1.5 Mini
Sub-130ms TTFB voice synthesis with 271+ voices across 15 languages, expressive prosody, and real-time SSE streaming for low-latency voice agents.

TTS 1.5 Max
Broadcast-quality voice synthesis with rich expressive prosody, 271+ voices across 15 languages, and real-time SSE streaming with per-word timestamps.
