> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.empiriolabs.ai/models/grok-imagine-image-2-0/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.empiriolabs.ai/_mcp/server. # Grok Imagine Image 2.0 > Text-to-image plus multi-reference editing with up to five source images, automatic or pinned quality tiers, and 1K or 2K output resolution. ![Grok Imagine Image 2.0](https://media.empiriolabs.ai/model-logos/grok-imagine-image-2-0.png) [xAI](/providers/xai) · Image Generation `POST /v1/images/generations` Text-to-image plus multi-reference editing with up to five source images, automatic or pinned quality tiers, and 1K or 2K output resolution. ## At a glance | Field | Value | | ------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- | | Model id | `grok-imagine-image-2-0` | | Model release date | 2026-08-07 | | Input modalities | Text, Image | | Output modalities | Image | | Context window | - | | Weight precision | - | | Features | image\_generation, image\_editing, multi\_image, text\_rendering | | Native inference | No | | New | Yes | | Supported endpoints | `POST /v1/images/generations`, `POST /v1/images/edits` | | Alternate model ids | `grok-imagine-image-2`, `grok-imagine-image-2.0`, `xai/grok-imagine-image-2-0`, `xai/grok-imagine-image-2`, `xai/grok-imagine-image-2.0` | ## Pricing | Charge | Spec | Rate | | ------------------ | --------- | ------- | | Low quality, 1K | per image | \$0.048 | | Low quality, 2K | per image | \$0.072 | | Medium quality, 1K | per image | \$0.072 | | Medium quality, 2K | per image | \$0.096 | | Image input | per image | \$0.05 | ## Example request ```bash curl https://api.empiriolabs.ai/v1/images/generations \ -H 'Authorization: Bearer $EMPIRIOLABS_API_KEY' \ -H 'Content-Type: application/json' \ -d '{"model": "grok-imagine-image-2-0", "prompt": "a misty forest at dawn", "size": "1024x1024"}' ``` ## Parameters | Parameter | Type | Required | Default | Description | | ----------------- | ------ | -------- | -------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `prompt` | string | no | - | Describe the image to create. With source images attached, this becomes the editing instruction. Refer to multiple images as \, \, and so on. Up to 8,000 characters. | | `image` | array | no | - | One to five source images to edit, as URLs, data URIs, or uploaded image objects. Leave empty to generate from the prompt alone. | | `images` | array | no | - | Plural alias for multiple uploaded image references. | | `quality` | enum | no | `"auto"` | Rendering quality. Auto lets the model choose the tier for each request, and billing follows the tier it serves. Low returns in a few seconds and costs less per image. Medium takes longer and resolves finer detail. · Allowed: `auto`, `low`, `medium` | | `resolution` | enum | no | `"1k"` | Output resolution tier. 2k renders more pixels and costs more per image than 1k. · Allowed: `1k`, `2k` | | `aspect_ratio` | enum | no | `"auto"` | Output aspect ratio. Auto lets the model pick the best fit for the prompt. For an edit, auto follows the first source image. Includes the 21:9 and 5:2 widescreen and banner formats. · Allowed: `auto`, `1:1`, `16:9`, `9:16`, `4:3`, `3:4`, `3:2`, `2:3`, `2:1`, `1:2`, `21:9`, `5:2`, `19.5:9`, `9:19.5`, `20:9`, `9:20` | | `num_images` | number | no | `1` | Number of output images. Each output image is billed separately. · Range: 1 – 10 | | `response_format` | enum | no | `"url"` | Return signed image URLs by default, or include base64 image data when b64\_json is requested. · Allowed: `url`, `b64_json` | ## Notes Generates images from a text prompt, or edits up to five source images from a natural-language instruction. **Modes** * Text-to-image: send a prompt with no image. * Image editing: attach one image and describe the change you want. * Multi-image editing: attach two to five images to combine subjects, transfer a style, or compose a scene. Refer to them in the prompt as \, \, and so on. **Example multi-image prompt** ``` Place the product from onto the marble surface in . ``` **Defaults** * Auto quality at 1K resolution * Auto aspect ratio, so the model picks the best fit for the prompt * One output image per request **Quality and resolution** Quality selects how much work goes into the render. Auto lets the model choose the tier for each request and currently serves low for a text prompt and medium for an edit. Low returns in a few seconds. Medium takes longer and resolves finer detail. Resolution selects the 1K or 2K output tier. Quality and resolution combine into four price points, so a low 1K image is the cheapest option and a medium 2K image the most detailed. **Controls** Supports prompt, source images (URL or upload), quality, resolution, aspect ratio across sixteen ratios including auto and the 21:9 and 5:2 widescreen formats, number of images, and response format. **Limits** * Up to 5 source images per edit * Up to 10 output images per request * Prompt up to 8,000 characters **Billing** * Charged per output image at the quality and resolution served, plus a per-image fee for each source image you attach. Text-to-image has no source-image fee. * With auto quality the charge follows the tier the model served for that request. Set quality to low or medium when you want the price fixed in advance. * A request that fails validation or is rejected before any image is produced is not billed. * An image that is produced but then blocked by content policy is still billed. This follows xAI usage policy, which charges for the generation even when the output is blocked by the content safety check. Keep prompts and source images within the usage policy to avoid this. **Content safety** Generated images pass an automated content-safety check. An image flagged as explicit is not returned. Because the image is generated before the check runs, a flagged generation is still billed. --- *Machine-readable schema:* `GET https://api.empiriolabs.ai/v1/models/grok-imagine-image-2-0`. > Text-to-image plus multi-reference editing with up to five source images, automatic or pinned quality tiers, and 1K or 2K output resolution.