StepAudio 3 TTS

POST /v1/audio/speechHuman-level speech synthesis across six languages, with 12 system voices, natural-language delivery direction, and low-latency streaming.
At a glance
Pricing
Example request
Parameters
Notes
Maximum input is 1,000 characters. Content in parentheses is treated as delivery direction and is not spoken. Use instruction for global emotion or style guidance, up to 500 characters. Twelve system voices are available, and a custom cloned voice ID is also accepted. Recognized languages are Chinese, English, Japanese, Korean, French, and Spanish, with languages other than Chinese and English in preview. Output formats are mp3, wav, flac, opus, and pcm at 8000, 16000, 22050, or 24000 Hz. Wav is returned as a complete file rather than streamed. This model does not support voice_label.
Machine-readable schema: GET https://api.empiriolabs.ai/v1/models/stepaudio-3-tts.
