AI Potluck
Back to Gap Map Model components / Multimodal models

GLM-V

Z.ai
open weights / Overall score: 3.7

GLM-V is Z.ai's vision-language reasoning line, from GLM-4.1V-Thinking through GLM-4.5V to GLM-4.6V and the 9B GLM-4.6V-Flash, trained with reinforcement learning on curriculum-sampled tasks. The models read long documents and video, ground objects with boxes, operate phone, desktop and web interfaces, and GLM-4.6V adds native function calling over images in a 128K context. The line is separate from Z.ai's GLM text models.

The older glm-4v-9b model, under the more restrictive GLM-4 license, is a predecessor and not part of this entry.

Openness

3 high confidence
3.0
weights
open(safetensors on the Hub, ungated)
data
described(the RLCS reinforcement-learning method is described
code
partial(inference, GUI-agent and desktop-assistant examples
license
MIT(OSI

The GLM-4.6V weights are published under the MIT license and the code repository under Apache 2.0. Z.ai describes its reinforcement-learning method but publishes no training data and no training code of its own.

Adoption

3 high confidence
3.0

Hugging Face downloads over the trailing 30 days, summed across Z.ai's GLM-4.1V, 4.5V and 4.6V checkpoints, including the FP8 builds. The general GLM text models the search also returns are excluded, as is the older glm-4v-9b.

Capability

4 high confidence
4.0

GLM-4.6V completes tasks on phones, desktops and the web, reads long documents and video, and calls tools with images as inputs. It is level with UI-TARS on what it can act on and, like it, takes no audio input.

Verified 2026-09-27