StepAudio 3 ASR Max

StepFun · Transcription
POST /v1/audio/transcriptionsContext-aware transcription for proper nouns, technical terms, dialects, and difficult audio including whispers, fast speech, and singing.
At a glance
Pricing
Example request
Parameters
Notes
Supports wav, mp3, ogg, m4a, and pcm input. PCM requests should include codec, sample rate, bit depth, and channel count. The language is detected automatically and cannot be set; Chinese, English, Japanese, Korean, French, and Spanish are recognized, with languages other than Chinese and English in preview. Inverse text normalization is on by default and can be turned off with enable_itn. Hotwords and word timestamps are not available on this model. Audio is billed by its measured duration, with a one second minimum.
Machine-readable schema: GET https://api.empiriolabs.ai/v1/models/stepaudio-3-asr-max.
