Seedance 1.5 Pro

ByteDance · Video Generation
POST /v1/videos/generationsJoint audio and video generation with millisecond lip sync, dialogue in six languages, cinematic camera moves, and up to 1080p output.
At a glance
Pricing
Example request
Parameters
Notes
ByteDance’s first joint audio-video model: sound (dialogue, effects, music) is generated together with the picture in one pass.
Modes
- Text to video, first frame, or first and last frame. No video, audio, or multi-image reference input.
- Put spoken lines in double quotes in the prompt for lip-synced dialogue. Speech covers English, Mandarin and several Chinese dialects, Japanese, Korean, Spanish, and Indonesian.
Output
- 480p, 720p, and 1080p at 24fps, 4 to 12 seconds (or let the model pick the length), MP4 with mono audio.
- generate_audio is on by default; silent output bills at the lower no-audio rate.
Draft mode
- draft renders a cheap 480p preview at a reduced rate. Re-render a pick at full quality by passing its job id as draft_task_id.
Input limits
- Images: jpeg, png, webp, bmp, tiff, gif, heic, or heif, 300 to 6000 px per side, aspect ratio between 0.4 and 2.5, under 30 MB each.
Billing
- Billed per second of generated video at the rate for the output resolution, with separate audio and silent rates. The model meters generated frames, so billed seconds track the rendered output; draft renders bill fewer seconds.
Machine-readable schema: GET https://api.empiriolabs.ai/v1/models/seedance-1-5-pro.
