Sonilo v1.1

Sonilo · Audio Generation
POST /v1/audio/generationsFrame-synced soundtracks generated straight from video, or music from a text prompt, with per-segment direction and stem separation.
At a glance
Pricing
Example request
Parameters
Notes
Modes
- auto scores the attached video, and falls back to text to music when there is no video
- Video mode follows the cuts and pacing of the source footage
- Text mode generates from a prompt alone
Controls
Supports prompt, mode, source video, 5 to 360 second duration in text mode, m4a, wav or mp3 output, 1 to 10 variants, stem separation, ducking, speech preservation, prompt influence, and JSON segment prompts.
Limits
- Source video up to 300MB and 6 minutes
- Duration 5 to 360 seconds in text mode. Video mode follows the length of the source
- Ducking, speech preservation and prompt influence apply to video mode only
Billing
- Charged per generated second at the catalog rate, with a 10 second minimum per track. A 4 second track bills 10 seconds.
- Each variant is billed separately and the 10 second minimum applies to each one.
- Stem separation, ducking and prompt influence add no charge.
- A request rejected before generation is not billed.
Machine-readable schema: GET https://api.empiriolabs.ai/v1/models/sonilo-v1-1.
