Seedance 2.5

ByteDance · Video Generation
POST /v1/videos/generationsLong-form video model for coherent clips up to 30 seconds, with up to 50 reference images, videos, and audio clips, native audio, editing, and extension.
At a glance
Pricing
Example request
Parameters
Notes
Next-generation Seedance built for long-form storytelling: one request can produce a coherent clip of up to 30 seconds, draw on up to 50 reference assets, and edit or extend existing footage.
Modes
- Text to video, first frame, or first and last frame.
- Multimodal reference to video from up to 30 images, 10 videos, and 10 audio clips. Audio-only input works on its own.
- To edit an attached clip or continue it, describe that in your prompt. The model reads the intent from your wording.
Output
- 480p and 720p at 24fps, 4 to 30 seconds, with native synchronized audio.
- MP4 by default. MOV carries higher color precision for grading and compositing, and some players cannot open it.
- When you supply a first frame, the output follows that image’s aspect ratio, so aspect_ratio stays adaptive. Reference jobs, including ones with a reference video, accept any ratio.
Input limits
- Images: jpeg, png, webp, bmp, tiff, gif, heic, or heif, 300 to 6000 px per side, under 30 MB each.
- Videos: MP4 or MOV, 2 to 30 seconds each and 30 seconds in total, under 200 MB each.
- Audio: WAV or MP3, 2 to 30 seconds each and 30 seconds in total, under 15 MB each.
- Reference uploads that contain real human faces are rejected by the model.
Billing
- Without a reference video, you are billed per second of generated video at the rate for the output resolution.
- With a reference video, the model prices the reference clip’s seconds alongside the generated seconds, at the lower video-input rate. Billed seconds are therefore reference duration plus output duration, so a request that includes a reference video costs more in total than the same output generated from text alone.
- Worked example at 480p: a 4 second clip from a text prompt bills 4 seconds. A 4 second clip generated from a 5 second reference video bills 9 seconds at the video-input rate.
Uploaded media preprocessing
- Video inputs are capped to 30 seconds.
- Uploaded video inputs are normalized to provider-compatible MP4 when needed. Set preprocess_video to false to send a clip through untouched.
Machine-readable schema: GET https://api.empiriolabs.ai/v1/models/seedance-2-5.
