> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.empiriolabs.ai/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.empiriolabs.ai/_mcp/server.

# Pixverse Avatar

> Turns a single portrait into a talking avatar video, driven by an uploaded audio track or by text spoken with a built-in voice.

![Pixverse Avatar](https://media.empiriolabs.ai/model-logos/pixverse-avatar.png)

[PixVerse](/providers/pixverse) · Video Generation

`POST /v1/videos/generations`

Turns a single portrait into a talking avatar video, driven by an uploaded audio track or by text spoken with a built-in voice.

## At a glance

| Field               | Value                                             |
| ------------------- | ------------------------------------------------- |
| Model id            | `pixverse-avatar`                                 |
| Model release date  | -                                                 |
| Input modalities    | Text, Image, Audio                                |
| Output modalities   | Video                                             |
| Context window      | -                                                 |
| Weight precision    | -                                                 |
| Features            | lipsync, image\_to\_video, audio\_in, audio\_sync |
| Native inference    | No                                                |
| New                 | No                                                |
| Supported endpoints | `POST /v1/videos/generations`                     |
| Alternate model ids | `pixverse/avatar`                                 |

## Pricing

| Charge | Spec       | Rate       |
| ------ | ---------- | ---------- |
| 360p   | per second | \$0.053333 |
| 540p   | per second | \$0.106667 |
| 720p   | per second | \$0.16     |
| 1080p  | per second | \$0.213333 |

## Example request

```bash
curl https://api.empiriolabs.ai/v1/videos/generations \
  -H 'Authorization: Bearer $EMPIRIOLABS_API_KEY' \
  -H 'Content-Type: application/json' \
  -d '{"model": "pixverse-avatar", "prompt": "sunrise over the ocean", "duration": 6}'
```

## Parameters

| Parameter                 | Type   | Required | Default  | Description                                                                                                                                                                                                                                                     |
| ------------------------- | ------ | -------- | -------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `lip_sync_tts_content`    | string | no       | -        | The line for the avatar to speak. Used when no audio file is attached. Billed in 15-character blocks.                                                                                                                                                           |
| `lip_sync_tts_speaker_id` | enum   | no       | `"Auto"` | Voice used for the spoken line. Auto picks one to suit the portrait. Ignored when an audio file is attached. · Allowed: `Auto`, `Emily`, `James`, `Isabella`, `Liam`, `Chloe`, `Adrian`, `Harper`, `Ava`, `Sophia`, `Julia`, `Mason`, `Jack`, `Oliver`, `Ethan` |
| `resolution`              | enum   | no       | `"720p"` | Output resolution. Higher resolutions bill at a higher per-second rate. · Allowed: `360p`, `540p`, `720p`, `1080p`                                                                                                                                              |
| `prompt`                  | string | no       | -        | Optional direction for delivery or framing.                                                                                                                                                                                                                     |
| `image`                   | string | no       | -        | Portrait image URL. Required.                                                                                                                                                                                                                                   |
| `audio`                   | string | no       | -        | Audio URL for the avatar to perform. Takes precedence over the spoken text.                                                                                                                                                                                     |

## Notes

Turns a single portrait image into a talking avatar video, driven either by an audio file you supply or by text spoken with a built-in voice.

**Inputs**

* `image`: the portrait to animate. Use a clear, front-facing subject.
* `audio`: the speech to perform, or `lip_sync_tts_content` with `lip_sync_tts_speaker_id` to have the model speak your text. Fourteen named voices are available plus `Auto`.
* `prompt` is optional and can describe delivery or framing.

**Output**

* 360p, 540p, 720p and 1080p.

**Billing**

* Billed per second of speech at the rate for the output resolution. Text-to-speech input is metered in 15-character blocks at the same rate, so the charge tracks the length of the script.

**Input limits**

* Image: PNG, JPEG or WebP, up to 20MB and 10000px on the long edge.
* Audio: MP3, WAV, M4A or AAC.

---

*Machine-readable schema:* `GET https://api.empiriolabs.ai/v1/models/pixverse-avatar`.