> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.empiriolabs.ai/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.empiriolabs.ai/_mcp/server.

# StepAudio 2.5 TTS

> Contextual StepFun text-to-speech model with natural-language voice direction and expressive delivery.

![StepAudio 2.5 TTS](https://media.empiriolabs.ai/model-logos/stepaudio-2-5-tts.png)

[StepFun](/providers/stepfun) · Audio Generation

`POST /v1/audio/speech`

Contextual StepFun text-to-speech model with natural-language voice direction and expressive delivery.

## At a glance

| Field               | Value                                                                         |
| ------------------- | ----------------------------------------------------------------------------- |
| Model id            | `stepaudio-2-5-tts`                                                           |
| Model release date  | 2026-04-21                                                                    |
| Input modalities    | Text                                                                          |
| Output modalities   | Audio                                                                         |
| Context window      | -                                                                             |
| Weight precision    | -                                                                             |
| Region              | International                                                                 |
| Features            | text\_to\_speech, contextual\_tts, voice\_control, multilingual               |
| Native inference    | No                                                                            |
| New                 | No                                                                            |
| Supported endpoints | `POST /v1/audio/speech`, `POST /v1/audio/speech:stream`                       |
| Alternate model ids | `stepaudio-2.5-tts`, `stepfun/stepaudio-2-5-tts`, `stepfun/stepaudio-2.5-tts` |

## Pricing

| Charge    | Spec                  | Rate   |
| --------- | --------------------- | ------ |
| Synthesis | per 10,000 characters | \$0.85 |

## Example request

```bash
curl https://api.empiriolabs.ai/v1/audio/speech \
  -H 'Authorization: Bearer $EMPIRIOLABS_API_KEY' \
  -H 'Content-Type: application/json' \
  -d '{"model": "stepaudio-2-5-tts", "input": "Hello from EmpirioLabs."}'
```

## Parameters

| Parameter           | Type    | Required | Default         | Description                                                                                                                                                                                                                               |
| ------------------- | ------- | -------- | --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `input`             | string  | yes      | -               | Text to synthesize. Maximum 1,000 characters. · Max: 1000                                                                                                                                                                                 |
| `voice`             | enum    | no       | `"lively-girl"` | Official StepFun voice. Custom cloned voice IDs are also accepted via the API. · Allowed: `lively-girl`, `vibrant-youth`, `soft-spoken-gentleman`, `magnetic-voiced-male`, `elegantgentle-female`, `livelybreezy-female`, `zixinnansheng` |
| `response_format`   | enum    | no       | `"mp3"`         | Output audio format. · Allowed: `mp3`, `wav`, `flac`, `opus`, `pcm`                                                                                                                                                                       |
| `speed`             | number  | no       | `1.0`           | Speech speed. · Range: 0.5 – 2                                                                                                                                                                                                            |
| `volume`            | number  | no       | `1.0`           | Output volume. · Range: 0.1 – 2                                                                                                                                                                                                           |
| `sample_rate`       | enum    | no       | `24000`         | Output sample rate in Hz. · Allowed: `8000`, `16000`, `22050`, `24000`                                                                                                                                                                    |
| `instruction`       | string  | no       | -               | Global emotion or style guidance. Only supported by StepAudio 2.5 TTS. · Max: 200                                                                                                                                                         |
| `pronunciation_map` | object  | no       | -               | Optional pronunciation mapping.                                                                                                                                                                                                           |
| `markdown_filter`   | boolean | no       | false           | Filter markdown syntax before synthesis.                                                                                                                                                                                                  |

## Notes

Maximum input is 1,000 characters. Content in parentheses is treated as direction and is not spoken. Use instruction for global emotion or style guidance. StepAudio 2.5 TTS does not support voice\_label.

---

*Machine-readable schema:* `GET https://api.empiriolabs.ai/v1/models/stepaudio-2-5-tts`.