Pixverse Avatar
PixVerse · Video Generation
POST /v1/videos/generationsTurns a single portrait into a talking avatar video, driven by an uploaded audio track or by text spoken with a built-in voice.
At a glance
Pricing
Example request
Parameters
Notes
Turns a single portrait image into a talking avatar video, driven either by an audio file you supply or by text spoken with a built-in voice.
Inputs
image: the portrait to animate. Use a clear, front-facing subject.audio: the speech to perform, orlip_sync_tts_contentwithlip_sync_tts_speaker_idto have the model speak your text. Fourteen named voices are available plusAuto.promptis optional and can describe delivery or framing.
Output
- 360p, 540p, 720p and 1080p.
Billing
- Billed per second of speech at the rate for the output resolution. Text-to-speech input is metered in 15-character blocks at the same rate, so the charge tracks the length of the script.
Input limits
- Image: PNG, JPEG or WebP, up to 20MB and 10000px on the long edge.
- Audio: MP3, WAV, M4A or AAC.
Machine-readable schema: GET https://api.empiriolabs.ai/v1/models/pixverse-avatar.
