> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.empiriolabs.ai/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.empiriolabs.ai/_mcp/server.

# GLM 4.6V Flash

> Free multimodal GLM-4.6V model for image, video, file, and text understanding with native function calling.

![GLM 4.6V Flash](https://media.empiriolabs.ai/model-logos/glm.png)

[Z.ai](/providers/zhipu) · Text Generation

`POST /v1/chat/completions`

Free multimodal GLM-4.6V model for image, video, file, and text understanding with native function calling.

## At a glance

| Field               | Value                                                                                                                        |
| ------------------- | ---------------------------------------------------------------------------------------------------------------------------- |
| Model id            | `glm-4-6v-flash`                                                                                                             |
| Model release date  | 2025-12-08                                                                                                                   |
| Input modalities    | Text, Image, Video, File                                                                                                     |
| Output modalities   | Text                                                                                                                         |
| Context window      | 128K                                                                                                                         |
| Weight precision    | -                                                                                                                            |
| Max output tokens   | 32,768                                                                                                                       |
| Region              | Singapore                                                                                                                    |
| Features            | vision, video\_understanding, document\_understanding, function\_calling, web\_search, reasoning                             |
| Native inference    | No                                                                                                                           |
| New                 | No                                                                                                                           |
| Structured output   | JSON Mode                                                                                                                    |
| Supported endpoints | `POST /v1/chat/completions`, `POST /v1/responses`, `POST /v1/messages`, `POST /v1beta/models/glm-4-6v-flash:generateContent` |
| Alternate model ids | `glm-4.6v-flash`, `zai/glm-4.6v-flash`, `zhipu/glm-4.6v-flash`                                                               |

## Pricing

| Charge              | Spec                       | Rate    |
| ------------------- | -------------------------- | ------- |
| Input               | per 1M prompt tokens       | Free    |
| Output              | per 1M generated tokens    | Free    |
| Implicit cache read | per 1M cached input tokens | Free    |
| Web search          | per request when enabled   | \$0.033 |

## Example request

```bash
curl https://api.empiriolabs.ai/v1/chat/completions \
  -H 'Authorization: Bearer $EMPIRIOLABS_API_KEY' \
  -H 'Content-Type: application/json' \
  -d '{"model": "glm-4-6v-flash", "messages": [{"role":"user","content":"Hello"}]}'
```

## Parameters

| Parameter               | Type    | Required | Default     | Description                                                                                                                                               |
| ----------------------- | ------- | -------- | ----------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `temperature`           | number  | no       | `1`         | Sampling temperature. Lower values are more deterministic. GLM-4.7-Flash and GLM-4.6V-Flash default to 1.0; GLM-4.5-Flash defaults to 0.6. · Range: 0 – 1 |
| `top_p`                 | number  | no       | `0.95`      | Nucleus sampling probability mass. Z.AI documents a 0.95 default for the GLM-4.7, GLM-4.6, and GLM-4.5 series. · Range: 0.01 – 1                          |
| `max_tokens`            | number  | no       | `4096`      | Maximum output tokens for GLM-4.6V-Flash: 32768. · Range: 1 – 32768                                                                                       |
| `stop`                  | array   | no       | -           | Stop word list. Z.AI currently supports one stop string in array form.                                                                                    |
| `do_sample`             | boolean | no       | true        | Enable sampling. When false, temperature and top\_p do not affect generation.                                                                             |
| `enable_thinking`       | boolean | no       | true        | Controls Z.AI thinking mode. Enabled is the default; GLM-4.6V-Flash automatically decides whether to think when enabled.                                  |
| `thinking`              | object  | no       | -           | Advanced thinking object. Use \{"type":"enabled"} or \{"type":"disabled"}. GLM-4.6V-Flash automatically decides whether to think when enabled.            |
| `tools`                 | array   | no       | -           | Function tools and the built-in web\_search tool are supported.                                                                                           |
| `tool_choice`           | enum    | no       | `"auto"`    | Controls whether the model may use tools. Z.AI documents auto tool selection; omit tools to disable tool use. · Allowed: `auto`                           |
| `tool_stream`           | boolean | no       | false       | Stream function-call tool output when stream is true. Z.AI documents tool\_stream for GLM-4.6 and newer models.                                           |
| `tool_web_search`       | boolean | no       | false       | Enable built-in web search. Adds \$0.033 per request when enabled.                                                                                        |
| `search_result`         | boolean | no       | true        | Return structured web search result metadata when web search is enabled.                                                                                  |
| `search_prompt`         | string  | no       | -           | Optional instruction for summarizing retrieved web search results.                                                                                        |
| `count`                 | number  | no       | `10`        | Number of web search results to retrieve. · Range: 1 – 50                                                                                                 |
| `search_domain_filter`  | string  | no       | -           | Optional domain whitelist for web search results.                                                                                                         |
| `search_recency_filter` | enum    | no       | `"noLimit"` | Optional web search recency window. · Allowed: `oneDay`, `oneWeek`, `oneMonth`, `oneYear`, `noLimit`                                                      |
| `response_format`       | enum    | no       | -           | Return the output as a valid JSON object (JSON mode). Describe the fields you want in your prompt.                                                        |

## Notes

Base token use is free. Built-in web search is optional through tool\_web\_search and adds \$0.033 per request when enabled.

---

*Machine-readable schema:* `GET https://api.empiriolabs.ai/v1/models/glm-4-6v-flash`.