GLM-V
Z.aiGLM-V is Z.ai's vision-language reasoning line, from GLM-4.1V-Thinking through GLM-4.5V to GLM-4.6V and the 9B GLM-4.6V-Flash, trained with reinforcement learning on curriculum-sampled tasks. The models read long documents and video, ground objects with boxes, operate phone, desktop and web interfaces, and GLM-4.6V adds native function calling over images in a 128K context. The line is separate from Z.ai's GLM text models.
The older glm-4v-9b model, under the more restrictive GLM-4 license, is a predecessor and not part of this entry.
Openness
3 high confidence- weights
- open(safetensors on the Hub, ungated)
- data
- described(the RLCS reinforcement-learning method is described
- code
- partial(inference, GUI-agent and desktop-assistant examples
- license
- MIT(OSI
The GLM-4.6V weights are published under the MIT license and the code repository under Apache 2.0. Z.ai describes its reinforcement-learning method but publishes no training data and no training code of its own.
- https://cdn.jsdelivr.net/gh/zai-org/GLM-V@main/LICENSE recorded 2026-09-27
The code repository LICENSE is the Apache License, Version 2.0
- https://cdn.jsdelivr.net/gh/zai-org/GLM-V@main/README.md recorded 2026-09-27
"LLaMA-Factory already supports fine-tuning for GLM-4.5V & GLM-4.1V-9B-Thinking models"; examples/gui-agent and examples/vlm-helper; RLCS named, no dataset linked
- https://huggingface.co/api/models?author=zai-org&search=GLM&limit=100 recorded 2026-09-27
The seven GLM-4.1V, 4.5V and 4.6V checkpoints are all tagged license:mit; glm-4v-9b is tagged other
- https://huggingface.co/api/models/zai-org/GLM-4.6V?expand[]=downloads&expand[]=cardData&expand[]=gated&expand[]=createdAt&expand[]=lastModified&expand[]=siblings recorded 2026-09-27
Hub metadata for GLM-4.6V: gated false, 41 safetensors shards, card license mit
Adoption
3 high confidenceHugging Face downloads over the trailing 30 days, summed across Z.ai's GLM-4.1V, 4.5V and 4.6V checkpoints, including the FP8 builds. The general GLM text models the search also returns are excluded, as is the older glm-4v-9b.
- https://huggingface.co/api/models?author=zai-org&search=GLM&limit=100 recorded 2026-09-27
Seven GLM-V checkpoints with 311,852 downloads in the trailing 30 days; GLM-4.1V-9B-Thinking 179,488 and GLM-4.6V-Flash 72,707
Capability
4 high confidenceGLM-4.6V completes tasks on phones, desktops and the web, reads long documents and video, and calls tools with images as inputs. It is level with UI-TARS on what it can act on and, like it, takes no audio input.
- https://arxiv.org/html/2507.01006 recorded 2026-09-27
GUI Agents rows: OSWorld 37.2, AndroidWorld 57.0, WebVoyager 81.0 for GLM-4.6V; Qwen2.5-VL-72B 8.8, 35.0, 40.4; MMLongBench-Doc 54.9
- https://huggingface.co/zai-org/GLM-4.6V/raw/main/README.md recorded 2026-09-27
"we integrate native Function Calling capabilities for the first time"; "up to 128K tokens of multi-document or long-document input"
Verified 2026-09-27