AI Potluck
Back to Gap Map Model components / Multimodal models

Kimi-VL

Moonshot AI
open weights / Overall score: 3.7

Kimi-VL is Moonshot AI's mixture-of-experts vision-language model, activating 2.8B parameters in its language decoder behind the native-resolution MoonViT encoder. It reads images, documents, multi-image sets and long video in a 128K context, and locates and operates interface elements on screen. The Thinking variants add long chain-of-thought reasoning; the line is separate from the general Kimi models.

Openness

3 high confidence
3.0
weights
open(safetensors on the Hub, ungated)
data
described(the long chain-of-thought SFT and RL stages are described
code
partial(inference and vLLM deployment only
license
MIT(OSI

Kimi-VL's weights are MIT, so the model can be used and modified freely. Moonshot publishes no training data and only inference code; fine-tuning is supported through a third-party framework.

Adoption

3 high confidence
3.0

Hugging Face downloads over the trailing 30 days, summed across Moonshot's three Kimi-VL checkpoints. Most of it is the non-thinking Instruct model.

Capability

4 medium confidence
4.0

Kimi-VL completes multi-step tasks in live Ubuntu and Windows environments, ahead of GPT-4o in its own report, on top of long-video, long-document and multi-image understanding. Its agent scores are low in absolute terms, but they place it level with UI-TARS rather than with the models that only locate interface elements.

Verified 2026-09-27