DeepSeek-VL
DeepSeekDeepSeek-VL is DeepSeek's line of vision-language models, released separately from its general chat models. DeepSeek-VL2 is a mixture-of-experts model in three sizes, from 1.0B to 4.5B activated parameters, that answers questions about images, reads documents, tables and charts, and grounds objects with bounding boxes across a few images per conversation.
DeepSeek has shipped no VL-named model since DeepSeek-VL2. Its later vision models, such as DeepSeek-V4-Flash-Vision-Exp, belong to the DeepSeek-V4 family and are part of the deepseek entry.
Openness
3 high confidence- weights
- open(safetensors for all three DeepSeek-VL2 sizes on the Hub, ungated)
- data
- described(the paper names open sources such as WIT, WikiHow and OBELICS alongside in-house knowledge, OCR and recaptioned data
- code
- partial(inference and demo code under MIT
- license
- DeepSeek-Model-License(the DeepSeek License Agreement 1.0
DeepSeek-VL2 may be used commercially under DeepSeek's model license, which forbids a list of harmful uses and requires passing those limits on. DeepSeek publishes inference code only, and much of the training data was built in house and not released.
- https://arxiv.org/html/2412.10302v1 recorded 2026-09-27
"we developed an in-house collection to expand coverage of general real-world knowledge"; "We combined these datasets with an extensive in-house OCR dataset"; training on "HAI-LLM", an internal platform
- https://cdn.jsdelivr.net/gh/deepseek-ai/DeepSeek-VL2@main/LICENSE-MODEL recorded 2026-09-27
DEEPSEEK LICENSE AGREEMENT Version 1.0; Attachment A: "You agree not to use the Model or Derivatives of the Model: ... For military use in any way"
- https://cdn.jsdelivr.net/gh/deepseek-ai/DeepSeek-VL2@main/README.md recorded 2026-09-27
"The use of DeepSeek-VL2 models is subject to DeepSeek Model License. DeepSeek-VL2 series supports commercial use."; install, inference.py and web demo only, no training section
- https://huggingface.co/api/models/deepseek-ai/deepseek-vl2?expand[]=downloads&expand[]=cardData&expand[]=gated&expand[]=createdAt&expand[]=lastModified&expand[]=siblings recorded 2026-09-27
Hub metadata for deepseek-vl2: gated false, safetensors shards, license other, license_name deepseek
Adoption
3 high confidenceHugging Face downloads over the trailing 30 days, summed across DeepSeek's DeepSeek-VL and DeepSeek-VL2 checkpoints. Most of it is the smallest DeepSeek-VL2 model.
- https://huggingface.co/api/models?author=deepseek-ai&search=vl&limit=1000&expand[]=downloads&expand[]=createdAt&expand[]=cardData&expand[]=gated&expand[]=lastModified recorded 2026-09-27
Seven DeepSeek-VL and DeepSeek-VL2 checkpoints with 119,734 downloads in the trailing 30 days; deepseek-vl2-tiny 90,047
Capability
2 high confidenceDeepSeek-VL2 reads documents, tables and charts with strong scores for its size and can point to objects in an image. Its paper says it handles only a few images per conversation and it has no video input, so it sits level with Moondream.
- https://arxiv.org/html/2412.10302v1 recorded 2026-09-27
Table 3, DeepSeek-VL2: DocVQA 93.3, ChartQA 86.0, InfoVQA 78.1, TextVQA 84.2, OCRBench 811; "DeepSeek-VL2's context window only allows for a few images per chat session"
- https://cdn.jsdelivr.net/gh/deepseek-ai/DeepSeek-VL2@main/README.md recorded 2026-09-27
"visual question answering, optical character recognition, document/table/chart understanding, and visual grounding"; a "multiple images/interleaved image-text" example
Verified 2026-09-27