> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.empiriolabs.ai/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.empiriolabs.ai/_mcp/server.

# Step TTS 2

> StepFun text-to-speech model with official voices, custom cloned voices, and voice tag controls.

![Step TTS 2](https://media.empiriolabs.ai/model-logos/step-tts-2.png)

[StepFun](/providers/stepfun) · Audio Generation

`POST /v1/audio/speech`

StepFun text-to-speech model with official voices, custom cloned voices, and voice tag controls.

## At a glance

| Field               | Value                                                   |
| ------------------- | ------------------------------------------------------- |
| Model id            | `step-tts-2`                                            |
| Model release date  | -                                                       |
| Input modalities    | Text                                                    |
| Output modalities   | Audio                                                   |
| Context window      | -                                                       |
| Weight precision    | -                                                       |
| Region              | International                                           |
| Features            | text\_to\_speech, voice\_tags, multilingual             |
| Native inference    | No                                                      |
| New                 | No                                                      |
| Supported endpoints | `POST /v1/audio/speech`, `POST /v1/audio/speech:stream` |
| Alternate model ids | `stepfun/step-tts-2`                                    |

## Pricing

| Charge    | Spec                  | Rate   |
| --------- | --------------------- | ------ |
| Synthesis | per 10,000 characters | \$0.40 |

## Example request

```bash
curl https://api.empiriolabs.ai/v1/audio/speech \
  -H 'Authorization: Bearer $EMPIRIOLABS_API_KEY' \
  -H 'Content-Type: application/json' \
  -d '{"model": "step-tts-2", "input": "Hello from EmpirioLabs."}'
```

## Parameters

| Parameter           | Type    | Required | Default         | Description                                                                                                                                                                                                                               |
| ------------------- | ------- | -------- | --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `input`             | string  | yes      | -               | Text to synthesize. Maximum 1,000 characters. · Max: 1000                                                                                                                                                                                 |
| `voice`             | enum    | no       | `"lively-girl"` | Official StepFun voice. Custom cloned voice IDs are also accepted via the API. · Allowed: `lively-girl`, `vibrant-youth`, `soft-spoken-gentleman`, `magnetic-voiced-male`, `elegantgentle-female`, `livelybreezy-female`, `zixinnansheng` |
| `response_format`   | enum    | no       | `"mp3"`         | Output audio format. · Allowed: `mp3`, `wav`, `flac`, `opus`, `pcm`                                                                                                                                                                       |
| `speed`             | number  | no       | `1.0`           | Speech speed. · Range: 0.5 – 2                                                                                                                                                                                                            |
| `volume`            | number  | no       | `1.0`           | Output volume. · Range: 0.1 – 2                                                                                                                                                                                                           |
| `sample_rate`       | enum    | no       | `24000`         | Output sample rate in Hz. · Allowed: `8000`, `16000`, `22050`, `24000`                                                                                                                                                                    |
| `voice_label`       | object  | no       | -               | Emotion and speaking-style tags for step-tts-2.                                                                                                                                                                                           |
| `pronunciation_map` | object  | no       | -               | Optional pronunciation mapping.                                                                                                                                                                                                           |
| `markdown_filter`   | boolean | no       | false           | Filter markdown syntax before synthesis.                                                                                                                                                                                                  |

## Notes

Maximum input is 1,000 characters. Supports official voice IDs, custom cloned voices, voice\_label emotion tags, and voice\_label style tags. The instruction parameter is only for stepaudio-2.5-tts.

---

*Machine-readable schema:* `GET https://api.empiriolabs.ai/v1/models/step-tts-2`.