Step 5 Preview

StepFun · Text Generation
POST /v1/chat/completionsFrontier multimodal reasoning with a 1M token context, image and video input, parallel tool calling, and strict JSON schema output.
At a glance
Pricing
Example request
Parameters
Notes
Supports text, image, and video input with a 1M token context, parallel function tools, strict JSON schema output, and reasoning_effort low, medium, or high. Prompt-cache hits are billed at the cache-read rate. Video input supports MP4 under 128 MB, with clips under 5 minutes recommended. This is a preview release, so its behavior and identifier may change.
Machine-readable schema: GET https://api.empiriolabs.ai/v1/models/step-5-preview.
