> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.empiriolabs.ai/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.empiriolabs.ai/_mcp/server.

# StepFun

> 17 StepFun models on EmpirioLabs, callable through one OpenAI-compatible API with shared billing and logs.

![StepFun](/_fern-img/86a65ac1241b0e73853026332da0415f93d5173793ce16f1df6262933b23c995.webp)

## StepFun

## Models from StepFun

![Step 5 Preview](https://media.empiriolabs.ai/model-logos/step-5-preview.png)

[Step 5 Preview](/models/step-5-preview)

Frontier multimodal reasoning with a 1M token context, image and video input, parallel tool calling, and strict JSON schema output.

![Step 3.7 Flash](https://media.empiriolabs.ai/model-logos/step-3-7-flash.png)

[Step 3.7 Flash](/models/step-3-7-flash)

StepFun multimodal reasoning model with image and video input, tool calling, adjustable reasoning effort, and 256K context.

![Step 3.5 Flash 2603](https://media.empiriolabs.ai/model-logos/step-3-5-flash-2603.png)

[Step 3.5 Flash 2603](/models/step-3-5-flash-2603)

Agent-optimized Step 3.5 Flash variant with low and high reasoning effort modes.

![Step 3.5 Flash](https://media.empiriolabs.ai/model-logos/step-3-5-flash.png)

[Step 3.5 Flash](/models/step-3-5-flash)

StepFun text reasoning model for agents, coding, tool calling, and long-context analysis.

![StepAudio 3 Chat](https://media.empiriolabs.ai/model-logos/stepaudio-3-chat.png)

[StepAudio 3 Chat](/models/stepaudio-3-chat)

Audio and text conversation model that reasons before answering, with function calling and JSON output.

![StepAudio 2.5 Chat](https://media.empiriolabs.ai/model-logos/stepaudio-2-5-chat.png)

[StepAudio 2.5 Chat](/models/stepaudio-2-5-chat)

StepFun audio and text conversation model with text output and paralinguistic understanding.

![StepAudio 3 TTS](https://media.empiriolabs.ai/model-logos/stepaudio-3-tts.png)

[StepAudio 3 TTS](/models/stepaudio-3-tts)

Human-level speech synthesis across six languages, with 12 system voices, natural-language delivery direction, and streaming playback.

![StepAudio 2.5 TTS](https://media.empiriolabs.ai/model-logos/stepaudio-2-5-tts.png)

[StepAudio 2.5 TTS](/models/stepaudio-2-5-tts)

Contextual StepFun text-to-speech model with natural-language voice direction and expressive delivery.

![StepAudio 3 Gen](https://media.empiriolabs.ai/model-logos/stepaudio-3-gen.png)

[StepAudio 3 Gen](/models/stepaudio-3-gen)

Generates dialogue, sound effects, ambience, and background music together in one clip from a written scene.

![StepAudio 3 Music](https://media.empiriolabs.ai/model-logos/stepaudio-3-music.png)

[StepAudio 3 Music](/models/stepaudio-3-music)

Writes complete songs from a style description, with vocals, arrangement, and mixing, or instrumentals and covers.

![StepAudio 3 ASR Max](https://media.empiriolabs.ai/model-logos/stepaudio-3-asr-max.png)

[StepAudio 3 ASR Max](/models/stepaudio-3-asr-max)

Context-aware transcription for proper nouns, technical terms, dialects, and difficult audio including whispers, fast speech, and singing.

![StepAudio 2.5 ASR Stream](https://media.empiriolabs.ai/model-logos/stepaudio-2-5-asr-stream.png)

[StepAudio 2.5 ASR Stream](/models/stepaudio-2-5-asr-stream)

Live transcription over a socket, returning partial results as the speaker talks, with voice activity detection marking each sentence.

![StepAudio 2.5 ASR](https://media.empiriolabs.ai/model-logos/stepaudio-2-5-asr.png)

[StepAudio 2.5 ASR](/models/stepaudio-2-5-asr)

StepFun streaming speech recognition model for Chinese and English audio transcription.

![StepAudio 3 Realtime](https://media.empiriolabs.ai/model-logos/stepaudio-3-realtime.png)

[StepAudio 3 Realtime](/models/stepaudio-3-realtime)

Speech in, speech out, over one socket. Interruptible mid-reply, with voice activity detection deciding when to answer.

![StepAudio 2.5 Realtime](https://media.empiriolabs.ai/model-logos/stepaudio-2-5-realtime.png)

[StepAudio 2.5 Realtime](/models/stepaudio-2-5-realtime)

Speech in, speech out, over one socket, with paralinguistic understanding of tone and pacing.

![Step Image Edit 2](https://media.empiriolabs.ai/model-logos/step-image-edit-2.png)

[Step Image Edit 2](/models/step-image-edit-2)

StepFun image generation and image editing model for text-to-image and single-image edits.

![Step TTS 2](https://media.empiriolabs.ai/model-logos/step-tts-2.png)

[Step TTS 2](/models/step-tts-2)

StepFun text-to-speech model with official voices, custom cloned voices, and voice tag controls.