Skip to navigation
StepFun

StepFun

Models from StepFun

Step 5 Preview
Step 5 Preview
Frontier multimodal reasoning with a 1M token context, image and video input, parallel tool calling, and strict JSON schema output.
Step 3.7 Flash
Step 3.7 Flash
StepFun multimodal reasoning model with image and video input, tool calling, adjustable reasoning effort, and 256K context.
Step 3.5 Flash 2603
Step 3.5 Flash 2603
Agent-optimized Step 3.5 Flash variant with low and high reasoning effort modes.
Step 3.5 Flash
Step 3.5 Flash
StepFun text reasoning model for agents, coding, tool calling, and long-context analysis.
StepAudio 3 Chat
StepAudio 3 Chat
Audio and text conversation model that reasons before answering, with function calling and JSON output.
StepAudio 2.5 Chat
StepAudio 2.5 Chat
StepFun audio and text conversation model with text output and paralinguistic understanding.
StepAudio 3 TTS
StepAudio 3 TTS
Human-level speech synthesis across six languages, with 12 system voices, natural-language delivery direction, and streaming playback.
StepAudio 2.5 TTS
StepAudio 2.5 TTS
Contextual StepFun text-to-speech model with natural-language voice direction and expressive delivery.
StepAudio 3 Gen
StepAudio 3 Gen
Generates dialogue, sound effects, ambience, and background music together in one clip from a written scene.
StepAudio 3 Music
StepAudio 3 Music
Writes complete songs from a style description, with vocals, arrangement, and mixing, or instrumentals and covers.
StepAudio 3 ASR Max
StepAudio 3 ASR Max
Context-aware transcription for proper nouns, technical terms, dialects, and difficult audio including whispers, fast speech, and singing.
StepAudio 2.5 ASR Stream
StepAudio 2.5 ASR Stream
Live transcription over a socket, returning partial results as the speaker talks, with voice activity detection marking each sentence.
StepAudio 2.5 ASR
StepAudio 2.5 ASR
StepFun streaming speech recognition model for Chinese and English audio transcription.
StepAudio 3 Realtime
StepAudio 3 Realtime
Speech in, speech out, over one socket. Interruptible mid-reply, with voice activity detection deciding when to answer.
StepAudio 2.5 Realtime
StepAudio 2.5 Realtime
Speech in, speech out, over one socket, with paralinguistic understanding of tone and pacing.
Step Image Edit 2
Step Image Edit 2
StepFun image generation and image editing model for text-to-image and single-image edits.
Step TTS 2
Step TTS 2
StepFun text-to-speech model with official voices, custom cloned voices, and voice tag controls.