Pixverse Lip Sync

PixVerse · Video Generation
POST /v1/videos/generationsAligns mouth movement in an existing video to an uploaded audio track or to text spoken by one of fourteen built-in voices.
At a glance
Pricing
Example request
Parameters
Notes
Aligns mouth movement in an existing video to speech, either from an audio file you supply or from text spoken by a built-in voice.
Inputs
video: the source clip, up to 300 seconds and 250MB.audio: the speech to match, up to 300 seconds and 250MB.- Or
lip_sync_tts_contentwithlip_sync_tts_speaker_idto have the model speak your text. Fourteen named voices are available plusAuto, which picks one to suit the clip.
Billing
- Billed per second of speech. Text-to-speech input is metered in 15-character blocks at the same rate, so the charge tracks the length of the script rather than the length of the source video.
Input limits
- Video: MP4, MOV or WebM, up to 1920px on the long edge.
- Audio: MP3, WAV, M4A or AAC.
Machine-readable schema: GET https://api.empiriolabs.ai/v1/models/pixverse-lipsync.
