> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.empiriolabs.ai/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.empiriolabs.ai/_mcp/server.

# TTS 2

> Realtime voice model that takes plain-English delivery direction, holds one voice identity across 200+ languages, and streams with word timestamps.

![TTS 2](https://media.empiriolabs.ai/model-logos/tts-2.png)

[Inworld](/providers/inworld) · Audio Generation

`POST /v1/audio/speech`

Realtime voice model that takes plain-English delivery direction, holds one voice identity across 200+ languages, and streams with word timestamps.

## At a glance

| Field               | Value                                                                                                                                             |
| ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- |
| Model id            | `tts-2`                                                                                                                                           |
| Model release date  | 2026-05-05                                                                                                                                        |
| Input modalities    | Text                                                                                                                                              |
| Output modalities   | Audio                                                                                                                                             |
| Context window      | -                                                                                                                                                 |
| Weight precision    | -                                                                                                                                                 |
| Features            | multi\_speaker, real\_time, low\_latency, streaming, word\_timestamps, character\_timestamps, multilingual, expressive\_prosody, voice\_direction |
| Native inference    | No                                                                                                                                                |
| New                 | No                                                                                                                                                |
| Supported endpoints | `POST /v1/audio/speech`, `POST /v1/audio/speech:stream`, `GET /v1/voices`                                                                         |
| Alternate model ids | `inworld-tts-2`                                                                                                                                   |

## Pricing

| Charge    | Spec              | Rate                  |
| --------- | ----------------- | --------------------- |
| Synthesis | per 1M characters | \$22.00 (was \$25.00) |

## Example request

```bash
curl https://api.empiriolabs.ai/v1/audio/speech \
  -H 'Authorization: Bearer $EMPIRIOLABS_API_KEY' \
  -H 'Content-Type: application/json' \
  -d '{"model": "tts-2", "input": "Hello from EmpirioLabs."}'
```

## Parameters

| Parameter                  | Type   | Required | Default   | Description                                                                                                                                                                                                                                                                                                                                                                                                                 |
| -------------------------- | ------ | -------- | --------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `input`                    | string | yes      | -         | Text to synthesize. Max 2,000 characters per request; chunk longer copy at sentence boundaries on the client. Optionally open the text with a bracketed delivery direction in plain English, for example `[whispering]` or `[Speak warmly, like you are greeting an old friend]`. The bracketed text steers delivery instead of being read aloud, and its characters count toward the billed character total. · Max: 2000   |
| `voice`                    | enum   | no       | `"Sarah"` | Voice preset. 20 hand-picked voices covering English, Spanish, Portuguese, Hindi, and various accents. For the full voice catalog (including cloned voices), use voice\_id instead. · Allowed: `Sarah`, `Olivia`, `Elizabeth`, `Ashley`, `Wendy`, `Julia`, `Priya`, `Pixie`, `Deborah`, `Alex`, `Mark`, `Edward`, `Theodore`, `Ronald`, `Dennis`, `Timothy`, `Shaun`, `Craig`, `Hades`, `Heitor`                            |
| `voice_id`                 | string | no       | -         | Free-form voice ID. Overrides voice when set. Use this to address any voice beyond the curated 20-preset list, including cloned and regional voices. Pass any voice name from GET /v1/voices. Example: Ashley, Mark.                                                                                                                                                                                                        |
| `language`                 | enum   | no       | `"en-US"` | BCP-47 language code. This model holds one voice identity across 200+ languages and locales, so you can pass any supported code here even if it is not in the dropdown. · Allowed: `en-US`, `en-GB`, `es-ES`, `es-MX`, `fr-FR`, `de-DE`, `it-IT`, `pt-BR`, `pt-PT`, `nl-NL`, `pl-PL`, `ru-RU`, `ja-JP`, `ko-KR`, `zh-CN`, `hi-IN`, `ar-EG`, `he-IL`, `tr-TR`, `vi-VN`, `th-TH`, `id-ID`, `uk-UA`, `el-GR`, `cs-CZ`, `sv-SE` |
| `output_format`            | enum   | no       | `"WAV"`   | Audio container/codec. WAV = LINEAR16 inside RIFF (ubiquitous). MP3 / OGG = compressed. PCM = headerless raw, useful for chunked real-time playback. FLAC = lossless. · Allowed: `MP3`, `WAV`, `OGG`, `FLAC`, `PCM`, `ALAW`, `MULAW`                                                                                                                                                                                        |
| `sample_rate`              | enum   | no       | `"24000"` | Output sample rate in Hz. 24000 is Inworld's default and what their voice models train at; raise to 48000 for broadcast quality. · Allowed: `8000`, `16000`, `22050`, `24000`, `32000`, `44100`, `48000`                                                                                                                                                                                                                    |
| `speed`                    | number | no       | `1.0`     | Speaking rate multiplier. 0.5 = half speed, 1.5 = 50% faster. · Range: 0.5 – 1.5                                                                                                                                                                                                                                                                                                                                            |
| `temperature`              | number | no       | `1.0`     | Voice expressiveness / variability. Lower = more consistent / "flat"; higher = more expressive but more variation between renders. · Range: 0.1 – 2.0                                                                                                                                                                                                                                                                       |
| `bit_rate`                 | number | no       | `128000`  | Bitrate in bps for MP3 / OGG\_OPUS. Ignored for other encodings. · Range: 32000 – 320000                                                                                                                                                                                                                                                                                                                                    |
| `apply_text_normalization` | enum   | no       | `"ON"`    | When ON, Inworld expands numbers / abbreviations / dates into spoken form ("USD 5" → "five US dollars"). · Allowed: `ON`, `OFF`                                                                                                                                                                                                                                                                                             |
| `timestamp_type`           | enum   | no       | `"NONE"`  | If non-NONE, the response includes per-word or per-character timestamps in timestamp\_info. Useful for caption / highlight UIs. · Allowed: `NONE`, `WORD`, `CHARACTER`                                                                                                                                                                                                                                                      |

## Notes

**Limits**

* Max input: 2,000 characters per request (chunk longer text at sentence boundaries)
* WebSocket: 20 concurrent connections, 5 contexts/connection
* Per-WS message: 1,000 characters

**Voice direction**

* Open the text with a bracketed instruction in plain English to steer delivery, for example `[whispering]`, `[excited]`, or `[Speak warmly, like you are greeting an old friend]`
* The bracketed text is treated as direction and is not read aloud
* Direction text counts toward the billed character total, so keep it short

**Latency**

* p90 TTFB: under 100 ms (Inworld benchmark)

**Voices**

* One voice identity held across 200+ languages and locales, including switching language mid-generation
* Curated presets in the dropdown; pass any other voice ID via `voice_id`

---

*Machine-readable schema:* `GET https://api.empiriolabs.ai/v1/models/tts-2`.