> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.empiriolabs.ai/models/skylark-embedding-vision/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.empiriolabs.ai/_mcp/server. # Skylark Embedding Vision > Multimodal embedding that fuses text, images, and video into one 1024 or 2048 dimension vector for cross-modal search and retrieval. ![Skylark Embedding Vision](https://media.empiriolabs.ai/model-logos/skylark-embedding-vision.png) [ByteDance](/providers/bytedance) · Embeddings `POST /v1/embeddings` Multimodal embedding that fuses text, images, and video into one 1024 or 2048 dimension vector for cross-modal search and retrieval. ## At a glance | Field | Value | | ------------------- | ------------------------------------------------------------------------------------------------- | | Model id | `skylark-embedding-vision` | | Model release date | 2025-06-28 | | Input modalities | Text, Image, Video | | Output modalities | Embedding | | Context window | 8K | | Weight precision | - | | Region | Malaysia | | Features | multimodal, fused vectors | | Native inference | No | | New | No | | Supported endpoints | `POST /v1/embeddings` | | Alternate model ids | `byteplus/skylark-embedding-vision`, `skylark-embedding-vision-250615`, `doubao-embedding-vision` | ## Pricing | Charge | Spec | Rate | | ------------------- | ------------- | ------ | | Text input | per 1M tokens | \$0.25 | | Image / video input | per 1M tokens | \$0.65 | ## Example request ```bash curl https://api.empiriolabs.ai/v1/embeddings \ -H 'Authorization: Bearer $EMPIRIOLABS_API_KEY' \ -H 'Content-Type: application/json' \ -d '{"model": "skylark-embedding-vision", "input": [{"type":"text","text":"Embed me."},{"type":"image","url":"https://media.empiriolabs.ai/example.jpg"}]}' ``` ## Parameters | Parameter | Type | Required | Default | Description | | ----------------- | ------ | -------- | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | | `input` | string | yes | - | Text to embed, an array of up to 16 strings (one embedding each), or an array of text, image\_url, and video\_url parts fused into one embedding. | | `dimensions` | enum | no | `"2048"` | Output vector dimensionality. · Allowed: `1024`, `2048` | | `encoding_format` | enum | no | `"float"` | Embedding encoding of the response. · Allowed: `float`, `base64` | | `instructions` | string | no | - | Optional retrieval instruction that conditions the embedding, such as a query-side or corpus-side template. | ## Notes **Output** * One fused vector per request across all input items (text, image, and video combine into a single embedding). Choose 1024 or 2048 dimensions. * An array of plain strings returns one embedding per string, up to 16 per request. **Per-input limits** * Text: up to 8,000 tokens per item. * Image: JPEG, PNG, WEBP, BMP, or TIFF, sides over 14 px, up to 36MP. * Video: MP4, AVI, or MOV, up to 50 MB; audio tracks are ignored. **Retrieval** * An optional instructions field conditions the embedding for the query or corpus side. Apply L2 normalization before cosine or dot-product comparison. --- *Machine-readable schema:* `GET https://api.empiriolabs.ai/v1/models/skylark-embedding-vision`. > Multimodal embedding that fuses text, images, and video into one 1024 or 2048 dimension vector for cross-modal search and retrieval.