AI Potluck
Back to Gap Map Model components / Multimodal models

InternVL

Shanghai AI Laboratory
open weights / Overall score: 4.0(strong)

InternVL is Shanghai AI Laboratory's open vision-language model family, released by OpenGVLab from InternVL 1.5 through InternVL3.5 in sizes from 1B to 241B. InternVL3.5 handles single and multi-image input, documents and video, and adds GUI interaction and embodied-agent tasks, with a Flash variant for faster inference. Its multimodal preference data and reinforcement-learning code are published alongside the weights.

Scored on InternVL3.5. The 78B checkpoints of InternVL2.5 and InternVL3 inherit the Qwen license from their Qwen2.5-72B base, while every InternVL3.5 size is Apache 2.0.

Openness

3 high confidence
3.0
weights
open(safetensors on the Hub, ungated)
data
partial(the MMPR-v1.2 and MMPR-Tiny preference and RL sets are released under MIT
code
open(training code and CascadeRL offline and online RL scripts (internvl_chat_gpt_oss))
license
Apache-2.0(OSI

Every InternVL3.5 checkpoint is Apache 2.0, and OpenGVLab publishes its training code and reinforcement-learning pipeline with the preference data that stage used. The supervised fine-tuning mixture for InternVL3.5 is not released, so the full model cannot be rebuilt.

Adoption

4 high confidence
4.0

Hugging Face downloads over the trailing 30 days, summed across OpenGVLab's InternVL checkpoints from InternVL 1.x to 3.5, including the transformers-format copies. Most of it is the small InternVL2 and InternVL3 models rather than the current generation.

Capability

4 high confidence
4.0

InternVL3.5 completes tasks in Windows and web environments as well as understanding documents, multi-image input and video. Its online agent scores sit beside UI-TARS-72B in its own report, which places it level with UI-TARS.

Verified 2026-09-27