Kimi-VL
Moonshot AIKimi-VL is Moonshot AI's mixture-of-experts vision-language model, activating 2.8B parameters in its language decoder behind the native-resolution MoonViT encoder. It reads images, documents, multi-image sets and long video in a 128K context, and locates and operates interface elements on screen. The Thinking variants add long chain-of-thought reasoning; the line is separate from the general Kimi models.
Openness
3 high confidence- weights
- open(safetensors on the Hub, ungated)
- data
- described(the long chain-of-thought SFT and RL stages are described
- code
- partial(inference and vLLM deployment only
- license
- MIT(OSI
Kimi-VL's weights are MIT, so the model can be used and modified freely. Moonshot publishes no training data and only inference code; fine-tuning is supported through a third-party framework.
- https://huggingface.co/api/datasets?author=moonshotai&limit=100 recorded 2026-09-27
Three moonshotai datasets, PerceptionBench, Kimi-Audio-GenTest and WorldVQA, all evaluation sets
- https://huggingface.co/api/models/moonshotai/Kimi-VL-A3B-Thinking-2506?expand[]=downloads&expand[]=cardData&expand[]=gated&expand[]=createdAt&expand[]=lastModified&expand[]=siblings recorded 2026-09-27
Hub metadata for Kimi-VL-A3B-Thinking-2506: gated false, seven safetensors shards, card license mit
- https://raw.githubusercontent.com/MoonshotAI/Kimi-VL/HEAD/LICENSE recorded 2026-09-27
The repository LICENSE is the MIT License, Copyright 2025 Moonshot AI
- https://raw.githubusercontent.com/MoonshotAI/Kimi-VL/HEAD/README.md recorded 2026-09-27
Thinking variant "Developed through long chain-of-thought (CoT) supervised fine-tuning (SFT) and reinforcement learning (RL)"; "Kimi-VL now offers seamless support for efficient fine-tuning through the latest version of LLaMA-Factory"
Adoption
3 high confidenceHugging Face downloads over the trailing 30 days, summed across Moonshot's three Kimi-VL checkpoints. Most of it is the non-thinking Instruct model.
- https://huggingface.co/api/models?author=moonshotai&search=Kimi-VL&sort=downloads&direction=-1&limit=100 recorded 2026-09-27
Three Kimi-VL checkpoints with 221,687 downloads in the trailing 30 days; Kimi-VL-A3B-Instruct 197,676
Capability
4 medium confidenceKimi-VL completes multi-step tasks in live Ubuntu and Windows environments, ahead of GPT-4o in its own report, on top of long-video, long-document and multi-image understanding. Its agent scores are low in absolute terms, but they place it level with UI-TARS rather than with the models that only locate interface elements.
- https://arxiv.org/pdf/2504.07491 recorded 2026-09-27
"For OSWorld, Kimi-VL reaches 8.22%, outperforming GPT-4o (5.03%) ... On WindowsAgentArena, our model achieves 10.4%"
- https://huggingface.co/moonshotai/Kimi-VL-A3B-Thinking-2506/raw/main/README.md recorded 2026-09-27
Performance table: MMMU 64.0, VideoMMMU 65.2, ScreenSpot-Pro 52.8, MMLongBench-DOC 42.1; "It Extends to Video Scenarios"
- https://raw.githubusercontent.com/MoonshotAI/Kimi-VL/HEAD/README.md recorded 2026-09-27
"Kimi-VL excels in multi-turn agent interaction tasks (e.g.,OSWorld)"; "Equipped with a 128K extended context window"
Verified 2026-09-27