AI Potluck
Back to Gap Map Model components / Multimodal models

DeepSeek-VL

DeepSeek
open weights / Overall score: 2.3

DeepSeek-VL is DeepSeek's line of vision-language models, released separately from its general chat models. DeepSeek-VL2 is a mixture-of-experts model in three sizes, from 1.0B to 4.5B activated parameters, that answers questions about images, reads documents, tables and charts, and grounds objects with bounding boxes across a few images per conversation.

DeepSeek has shipped no VL-named model since DeepSeek-VL2. Its later vision models, such as DeepSeek-V4-Flash-Vision-Exp, belong to the DeepSeek-V4 family and are part of the deepseek entry.

Openness

3 high confidence
3.0
weights
open(safetensors for all three DeepSeek-VL2 sizes on the Hub, ungated)
data
described(the paper names open sources such as WIT, WikiHow and OBELICS alongside in-house knowledge, OCR and recaptioned data
code
partial(inference and demo code under MIT
license
DeepSeek-Model-License(the DeepSeek License Agreement 1.0

DeepSeek-VL2 may be used commercially under DeepSeek's model license, which forbids a list of harmful uses and requires passing those limits on. DeepSeek publishes inference code only, and much of the training data was built in house and not released.

Adoption

3 high confidence
3.0

Hugging Face downloads over the trailing 30 days, summed across DeepSeek's DeepSeek-VL and DeepSeek-VL2 checkpoints. Most of it is the smallest DeepSeek-VL2 model.

Capability

2 high confidence
2.0

DeepSeek-VL2 reads documents, tables and charts with strong scores for its size and can point to objects in an image. Its paper says it handles only a few images per conversation and it has no video input, so it sits level with Moondream.

Verified 2026-09-27