> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.empiriolabs.ai/models/pixverse-avatar/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.empiriolabs.ai/_mcp/server. # Pixverse Avatar > Turns a single portrait into a talking avatar video, driven by an uploaded audio track or by text spoken with a built-in voice. ![Pixverse Avatar](https://media.empiriolabs.ai/model-logos/pixverse-avatar.png) [PixVerse](/providers/pixverse) · Video Generation `POST /v1/videos/generations` Turns a single portrait into a talking avatar video, driven by an uploaded audio track or by text spoken with a built-in voice. ## At a glance | Field | Value | | ------------------- | ------------------------------------------------- | | Model id | `pixverse-avatar` | | Model release date | - | | Input modalities | Text, Image, Audio | | Output modalities | Video | | Context window | - | | Weight precision | - | | Features | lipsync, image\_to\_video, audio\_in, audio\_sync | | Native inference | No | | New | No | | Supported endpoints | `POST /v1/videos/generations` | | Alternate model ids | `pixverse/avatar` | ## Pricing | Charge | Spec | Rate | | ------ | ---------- | ---------- | | 360p | per second | \$0.053333 | | 540p | per second | \$0.106667 | | 720p | per second | \$0.16 | | 1080p | per second | \$0.213333 | ## Example request ```bash curl https://api.empiriolabs.ai/v1/videos/generations \ -H 'Authorization: Bearer $EMPIRIOLABS_API_KEY' \ -H 'Content-Type: application/json' \ -d '{"model": "pixverse-avatar", "prompt": "sunrise over the ocean", "duration": 6}' ``` ## Parameters | Parameter | Type | Required | Default | Description | | ------------------------- | ------ | -------- | -------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `lip_sync_tts_content` | string | no | - | The line for the avatar to speak. Used when no audio file is attached. Billed in 15-character blocks. | | `lip_sync_tts_speaker_id` | enum | no | `"Auto"` | Voice used for the spoken line. Auto picks one to suit the portrait. Ignored when an audio file is attached. · Allowed: `Auto`, `Emily`, `James`, `Isabella`, `Liam`, `Chloe`, `Adrian`, `Harper`, `Ava`, `Sophia`, `Julia`, `Mason`, `Jack`, `Oliver`, `Ethan` | | `resolution` | enum | no | `"720p"` | Output resolution. Higher resolutions bill at a higher per-second rate. · Allowed: `360p`, `540p`, `720p`, `1080p` | | `prompt` | string | no | - | Optional direction for delivery or framing. | | `image` | string | no | - | Portrait image URL. Required. | | `audio` | string | no | - | Audio URL for the avatar to perform. Takes precedence over the spoken text. | ## Notes Turns a single portrait image into a talking avatar video, driven either by an audio file you supply or by text spoken with a built-in voice. **Inputs** * `image`: the portrait to animate. Use a clear, front-facing subject. * `audio`: the speech to perform, or `lip_sync_tts_content` with `lip_sync_tts_speaker_id` to have the model speak your text. Fourteen named voices are available plus `Auto`. * `prompt` is optional and can describe delivery or framing. **Output** * 360p, 540p, 720p and 1080p. **Billing** * Billed per second of speech at the rate for the output resolution. Text-to-speech input is metered in 15-character blocks at the same rate, so the charge tracks the length of the script. **Input limits** * Image: PNG, JPEG or WebP, up to 20MB and 10000px on the long edge. * Audio: MP3, WAV, M4A or AAC. --- *Machine-readable schema:* `GET https://api.empiriolabs.ai/v1/models/pixverse-avatar`. > Turns a single portrait into a talking avatar video, driven by an uploaded audio track or by text spoken with a built-in voice.