> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.empiriolabs.ai/models/glm-4-6v-flash/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.empiriolabs.ai/_mcp/server. # GLM 4.6V Flash > Free multimodal GLM-4.6V model for image, video, file, and text understanding with native function calling. ![GLM 4.6V Flash](https://media.empiriolabs.ai/model-logos/glm.png) [Z.ai](/providers/zhipu) · Text Generation `POST /v1/chat/completions` Free multimodal GLM-4.6V model for image, video, file, and text understanding with native function calling. ## At a glance | Field | Value | | ------------------- | ---------------------------------------------------------------------------------------------------------------------------- | | Model id | `glm-4-6v-flash` | | Model release date | 2025-12-08 | | Input modalities | Text, Image, Video, File | | Output modalities | Text | | Context window | 128K | | Weight precision | - | | Max output tokens | 32,768 | | Region | Singapore | | Features | vision, video\_understanding, document\_understanding, function\_calling, web\_search, reasoning | | Native inference | No | | New | No | | Structured output | JSON Mode | | Supported endpoints | `POST /v1/chat/completions`, `POST /v1/responses`, `POST /v1/messages`, `POST /v1beta/models/glm-4-6v-flash:generateContent` | | Alternate model ids | `glm-4.6v-flash`, `zai/glm-4.6v-flash`, `zhipu/glm-4.6v-flash` | ## Pricing | Charge | Spec | Rate | | ------------------- | -------------------------- | ------- | | Input | per 1M prompt tokens | Free | | Output | per 1M generated tokens | Free | | Implicit cache read | per 1M cached input tokens | Free | | Web search | per request when enabled | \$0.033 | ## Example request ```bash curl https://api.empiriolabs.ai/v1/chat/completions \ -H 'Authorization: Bearer $EMPIRIOLABS_API_KEY' \ -H 'Content-Type: application/json' \ -d '{"model": "glm-4-6v-flash", "messages": [{"role":"user","content":"Hello"}]}' ``` ## Parameters | Parameter | Type | Required | Default | Description | | ----------------------- | ------- | -------- | ----------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- | | `temperature` | number | no | `1` | Sampling temperature. Lower values are more deterministic. GLM-4.7-Flash and GLM-4.6V-Flash default to 1.0; GLM-4.5-Flash defaults to 0.6. · Range: 0 – 1 | | `top_p` | number | no | `0.95` | Nucleus sampling probability mass. Z.AI documents a 0.95 default for the GLM-4.7, GLM-4.6, and GLM-4.5 series. · Range: 0.01 – 1 | | `max_tokens` | number | no | `4096` | Maximum output tokens for GLM-4.6V-Flash: 32768. · Range: 1 – 32768 | | `stop` | array | no | - | Stop word list. Z.AI currently supports one stop string in array form. | | `do_sample` | boolean | no | true | Enable sampling. When false, temperature and top\_p do not affect generation. | | `enable_thinking` | boolean | no | true | Controls Z.AI thinking mode. Enabled is the default; GLM-4.6V-Flash automatically decides whether to think when enabled. | | `thinking` | object | no | - | Advanced thinking object. Use \{"type":"enabled"} or \{"type":"disabled"}. GLM-4.6V-Flash automatically decides whether to think when enabled. | | `tools` | array | no | - | Function tools and the built-in web\_search tool are supported. | | `tool_choice` | enum | no | `"auto"` | Controls whether the model may use tools. Z.AI documents auto tool selection; omit tools to disable tool use. · Allowed: `auto` | | `tool_stream` | boolean | no | false | Stream function-call tool output when stream is true. Z.AI documents tool\_stream for GLM-4.6 and newer models. | | `tool_web_search` | boolean | no | false | Enable built-in web search. Adds \$0.033 per request when enabled. | | `search_result` | boolean | no | true | Return structured web search result metadata when web search is enabled. | | `search_prompt` | string | no | - | Optional instruction for summarizing retrieved web search results. | | `count` | number | no | `10` | Number of web search results to retrieve. · Range: 1 – 50 | | `search_domain_filter` | string | no | - | Optional domain whitelist for web search results. | | `search_recency_filter` | enum | no | `"noLimit"` | Optional web search recency window. · Allowed: `oneDay`, `oneWeek`, `oneMonth`, `oneYear`, `noLimit` | | `response_format` | enum | no | - | Return the output as a valid JSON object (JSON mode). Describe the fields you want in your prompt. | ## Notes Base token use is free. Built-in web search is optional through tool\_web\_search and adds \$0.033 per request when enabled. --- *Machine-readable schema:* `GET https://api.empiriolabs.ai/v1/models/glm-4-6v-flash`. > Free multimodal GLM-4.6V model for image, video, file, and text understanding with native function calling.