> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.empiriolabs.ai/models/step-tts-2/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.empiriolabs.ai/_mcp/server. # Step TTS 2 > StepFun text-to-speech model with official voices, custom cloned voices, and voice tag controls. ![Step TTS 2](https://media.empiriolabs.ai/model-logos/step-tts-2.png) [StepFun](/providers/stepfun) · Audio Generation `POST /v1/audio/speech` StepFun text-to-speech model with official voices, custom cloned voices, and voice tag controls. ## At a glance | Field | Value | | ------------------- | ------------------------------------------------------- | | Model id | `step-tts-2` | | Model release date | - | | Input modalities | Text | | Output modalities | Audio | | Context window | - | | Weight precision | - | | Region | International | | Features | text\_to\_speech, voice\_tags, multilingual | | Native inference | No | | New | No | | Supported endpoints | `POST /v1/audio/speech`, `POST /v1/audio/speech:stream` | | Alternate model ids | `stepfun/step-tts-2` | ## Pricing | Charge | Spec | Rate | | --------- | --------------------- | ------ | | Synthesis | per 10,000 characters | \$0.40 | ## Example request ```bash curl https://api.empiriolabs.ai/v1/audio/speech \ -H 'Authorization: Bearer $EMPIRIOLABS_API_KEY' \ -H 'Content-Type: application/json' \ -d '{"model": "step-tts-2", "input": "Hello from EmpirioLabs."}' ``` ## Parameters | Parameter | Type | Required | Default | Description | | ------------------- | ------- | -------- | --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `input` | string | yes | - | Text to synthesize. Maximum 1,000 characters. · Max: 1000 | | `voice` | enum | no | `"lively-girl"` | Official StepFun voice. Custom cloned voice IDs are also accepted via the API. · Allowed: `lively-girl`, `vibrant-youth`, `soft-spoken-gentleman`, `magnetic-voiced-male`, `elegantgentle-female`, `livelybreezy-female`, `zixinnansheng` | | `response_format` | enum | no | `"mp3"` | Output audio format. · Allowed: `mp3`, `wav`, `flac`, `opus`, `pcm` | | `speed` | number | no | `1.0` | Speech speed. · Range: 0.5 – 2 | | `volume` | number | no | `1.0` | Output volume. · Range: 0.1 – 2 | | `sample_rate` | enum | no | `24000` | Output sample rate in Hz. · Allowed: `8000`, `16000`, `22050`, `24000` | | `voice_label` | object | no | - | Emotion and speaking-style tags for step-tts-2. | | `pronunciation_map` | object | no | - | Optional pronunciation mapping. | | `markdown_filter` | boolean | no | false | Filter markdown syntax before synthesis. | ## Notes Maximum input is 1,000 characters. Supports official voice IDs, custom cloned voices, voice\_label emotion tags, and voice\_label style tags. The instruction parameter is only for stepaudio-2.5-tts. --- *Machine-readable schema:* `GET https://api.empiriolabs.ai/v1/models/step-tts-2`. > StepFun text-to-speech model with official voices, custom cloned voices, and voice tag controls.