StepAudio 3 Gen

StepFun · Audio Generation
POST /v1/audio/generationsGenerates dialogue, sound effects, ambience, and background music together in one clip from a written scene.
At a glance
Pricing
Example request
Parameters
Notes
Describe a scene and the model performs it as one finished clip: spoken lines, sound effects, ambience, and background music together. Send a prompt, or use roles and scripts for multi-speaker control. Wrap a sound effect or music cue in square brackets and a delivery direction in parentheses. Roles and instruction are capped at 500 characters each and scripts at 1,000. Output formats are mp3, wav, flac, opus, and pcm. Reference voices and voice cloning are not available on this model. This model is in preview and its capabilities may change.
Machine-readable schema: GET https://api.empiriolabs.ai/v1/models/stepaudio-3-gen.
