> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.empiriolabs.ai/providers/stepfun/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.empiriolabs.ai/_mcp/server. # StepFun > 17 StepFun models on EmpirioLabs, callable through one OpenAI-compatible API with shared billing and logs. ![StepFun](/_fern-img/86a65ac1241b0e73853026332da0415f93d5173793ce16f1df6262933b23c995.webp) ## StepFun ## Models from StepFun ![Step 5 Preview](https://media.empiriolabs.ai/model-logos/step-5-preview.png) [Step 5 Preview](/models/step-5-preview) Frontier multimodal reasoning with a 1M token context, image and video input, parallel tool calling, and strict JSON schema output. ![Step 3.7 Flash](https://media.empiriolabs.ai/model-logos/step-3-7-flash.png) [Step 3.7 Flash](/models/step-3-7-flash) StepFun multimodal reasoning model with image and video input, tool calling, adjustable reasoning effort, and 256K context. ![Step 3.5 Flash 2603](https://media.empiriolabs.ai/model-logos/step-3-5-flash-2603.png) [Step 3.5 Flash 2603](/models/step-3-5-flash-2603) Agent-optimized Step 3.5 Flash variant with low and high reasoning effort modes. ![Step 3.5 Flash](https://media.empiriolabs.ai/model-logos/step-3-5-flash.png) [Step 3.5 Flash](/models/step-3-5-flash) StepFun text reasoning model for agents, coding, tool calling, and long-context analysis. ![StepAudio 3 Chat](https://media.empiriolabs.ai/model-logos/stepaudio-3-chat.png) [StepAudio 3 Chat](/models/stepaudio-3-chat) Audio and text conversation model that reasons before answering, with function calling and JSON output. ![StepAudio 2.5 Chat](https://media.empiriolabs.ai/model-logos/stepaudio-2-5-chat.png) [StepAudio 2.5 Chat](/models/stepaudio-2-5-chat) StepFun audio and text conversation model with text output and paralinguistic understanding. ![StepAudio 3 TTS](https://media.empiriolabs.ai/model-logos/stepaudio-3-tts.png) [StepAudio 3 TTS](/models/stepaudio-3-tts) Human-level speech synthesis across six languages, with 12 system voices, natural-language delivery direction, and streaming playback. ![StepAudio 2.5 TTS](https://media.empiriolabs.ai/model-logos/stepaudio-2-5-tts.png) [StepAudio 2.5 TTS](/models/stepaudio-2-5-tts) Contextual StepFun text-to-speech model with natural-language voice direction and expressive delivery. ![StepAudio 3 Gen](https://media.empiriolabs.ai/model-logos/stepaudio-3-gen.png) [StepAudio 3 Gen](/models/stepaudio-3-gen) Generates dialogue, sound effects, ambience, and background music together in one clip from a written scene. ![StepAudio 3 Music](https://media.empiriolabs.ai/model-logos/stepaudio-3-music.png) [StepAudio 3 Music](/models/stepaudio-3-music) Writes complete songs from a style description, with vocals, arrangement, and mixing, or instrumentals and covers. ![StepAudio 3 ASR Max](https://media.empiriolabs.ai/model-logos/stepaudio-3-asr-max.png) [StepAudio 3 ASR Max](/models/stepaudio-3-asr-max) Context-aware transcription for proper nouns, technical terms, dialects, and difficult audio including whispers, fast speech, and singing. ![StepAudio 2.5 ASR Stream](https://media.empiriolabs.ai/model-logos/stepaudio-2-5-asr-stream.png) [StepAudio 2.5 ASR Stream](/models/stepaudio-2-5-asr-stream) Live transcription over a socket, returning partial results as the speaker talks, with voice activity detection marking each sentence. ![StepAudio 2.5 ASR](https://media.empiriolabs.ai/model-logos/stepaudio-2-5-asr.png) [StepAudio 2.5 ASR](/models/stepaudio-2-5-asr) StepFun streaming speech recognition model for Chinese and English audio transcription. ![StepAudio 3 Realtime](https://media.empiriolabs.ai/model-logos/stepaudio-3-realtime.png) [StepAudio 3 Realtime](/models/stepaudio-3-realtime) Speech in, speech out, over one socket. Interruptible mid-reply, with voice activity detection deciding when to answer. ![StepAudio 2.5 Realtime](https://media.empiriolabs.ai/model-logos/stepaudio-2-5-realtime.png) [StepAudio 2.5 Realtime](/models/stepaudio-2-5-realtime) Speech in, speech out, over one socket, with paralinguistic understanding of tone and pacing. ![Step Image Edit 2](https://media.empiriolabs.ai/model-logos/step-image-edit-2.png) [Step Image Edit 2](/models/step-image-edit-2) StepFun image generation and image editing model for text-to-image and single-image edits. ![Step TTS 2](https://media.empiriolabs.ai/model-logos/step-tts-2.png) [Step TTS 2](/models/step-tts-2) StepFun text-to-speech model with official voices, custom cloned voices, and voice tag controls. > 17 StepFun models on EmpirioLabs, callable through one OpenAI-compatible API with shared billing and logs.