> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.empiriolabs.ai/models/stepaudio-2-5-asr/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.empiriolabs.ai/_mcp/server. # StepAudio 2.5 ASR > StepFun streaming speech recognition model for Chinese and English audio transcription. ![StepAudio 2.5 ASR](https://media.empiriolabs.ai/model-logos/stepaudio-2-5-asr.png) [StepFun](/providers/stepfun) · Transcription `POST /v1/audio/transcriptions` StepFun streaming speech recognition model for Chinese and English audio transcription. ## At a glance | Field | Value | | ------------------- | ----------------------------------------------------------------------------- | | Model id | `stepaudio-2-5-asr` | | Model release date | 2026-04-24 | | Input modalities | Audio | | Output modalities | Text | | Context window | - | | Weight precision | - | | Region | International | | Features | transcription, speech\_to\_text, streaming\_asr, timestamps | | Native inference | No | | New | No | | Supported endpoints | `POST /v1/audio/transcriptions` | | Alternate model ids | `stepaudio-2.5-asr`, `stepfun/stepaudio-2-5-asr`, `stepfun/stepaudio-2.5-asr` | ## Pricing | Charge | Spec | Rate | | ------------- | ----------------- | ------- | | Transcription | per hour of audio | \$0.022 | ## Example request ```bash curl https://api.empiriolabs.ai/v1/audio/transcriptions \ -H 'Authorization: Bearer $EMPIRIOLABS_API_KEY' \ -F model=stepaudio-2-5-asr \ -F file=@meeting.mp3 ``` ## Parameters | Parameter | Type | Required | Default | Description | | ------------------ | ------- | -------- | ------- | --------------------------------------------------------------- | | `file` | audio | no | - | Audio file upload for transcription. | | `file_url` | string | no | - | Public URL to an audio file. | | `audio_base64` | string | no | - | Base64 audio payload for JSON requests. | | `language` | enum | no | `"en"` | Recognition language. · Allowed: `en`, `zh` | | `hotwords` | array | no | - | Optional hotword list. | | `enable_itn` | boolean | no | true | Enable inverse text normalization. | | `enable_timestamp` | boolean | no | false | Return timestamp fields from StepFun SSE events when available. | | `format` | enum | no | `"wav"` | Audio container format. · Allowed: `wav`, `mp3`, `ogg`, `pcm` | | `codec` | string | no | - | PCM codec such as pcm\_s16le. | | `rate` | integer | no | `16000` | PCM sample rate in Hz. · Range: 8000 – 48000 | | `bits` | integer | no | `16` | PCM bit depth. · Range: 8 – 32 | | `channel` | integer | no | `1` | PCM channel count. · Range: 1 – 8 | ## Notes Supports wav, mp3, ogg, and pcm input. PCM requests should include codec, sample rate, bit depth, and channel count. Recognition language supports Chinese and English. Timestamp output is available through enable\_timestamp. --- *Machine-readable schema:* `GET https://api.empiriolabs.ai/v1/models/stepaudio-2-5-asr`. > StepFun streaming speech recognition model for Chinese and English audio transcription.