InternVL
Shanghai AI LaboratoryInternVL is Shanghai AI Laboratory's open vision-language model family, released by OpenGVLab from InternVL 1.5 through InternVL3.5 in sizes from 1B to 241B. InternVL3.5 handles single and multi-image input, documents and video, and adds GUI interaction and embodied-agent tasks, with a Flash variant for faster inference. Its multimodal preference data and reinforcement-learning code are published alongside the weights.
Scored on InternVL3.5. The 78B checkpoints of InternVL2.5 and InternVL3 inherit the Qwen license from their Qwen2.5-72B base, while every InternVL3.5 size is Apache 2.0.
Openness
3 high confidence- weights
- open(safetensors on the Hub, ungated)
- data
- partial(the MMPR-v1.2 and MMPR-Tiny preference and RL sets are released under MIT
- code
- open(training code and CascadeRL offline and online RL scripts (internvl_chat_gpt_oss))
- license
- Apache-2.0(OSI
Every InternVL3.5 checkpoint is Apache 2.0, and OpenGVLab publishes its training code and reinforcement-learning pipeline with the preference data that stage used. The supervised fine-tuning mixture for InternVL3.5 is not released, so the full model cannot be rebuilt.
- https://cdn.jsdelivr.net/gh/OpenGVLab/InternVL@main/internvl_chat_gpt_oss/shell/internvl3_5_gpt_oss/internvl3_5_gpt_oss_20b_stage3_mpo.sh recorded 2026-09-27
The stage-3 MPO training launch script for InternVL3.5-GPT-OSS-20B-A4B
- https://cdn.jsdelivr.net/gh/OpenGVLab/InternVL@main/README.md recorded 2026-09-27
"We open-source the training code of InternVL3_5-GPT-OSS-20B-A4B and CascadeRL ... The training data for these two stages (MMPR-v1.2 and MMPR-Tiny) are also open-sourced."
- https://huggingface.co/api/datasets/OpenGVLab/MMPR-v1.2 recorded 2026-09-27
MMPR-v1.2: public, ungated, MIT, preference data that "greatly improves the overall performance of InternVL3.5"
- https://huggingface.co/api/models?author=OpenGVLab&search=InternVL3_5&limit=100 recorded 2026-09-27
50 InternVL3.5 checkpoints from 1B to 241B-A28B, every one tagged license:apache-2.0
- https://huggingface.co/api/models/OpenGVLab/InternVL3_5-8B?expand[]=downloads&expand[]=cardData&expand[]=gated&expand[]=createdAt&expand[]=lastModified&expand[]=siblings recorded 2026-09-27
Hub metadata for InternVL3_5-8B: gated false, safetensors, card license apache-2.0, datasets OpenGVLab/MMPR-v1.2 and MMPR-Tiny
Adoption
4 high confidenceHugging Face downloads over the trailing 30 days, summed across OpenGVLab's InternVL checkpoints from InternVL 1.x to 3.5, including the transformers-format copies. Most of it is the small InternVL2 and InternVL3 models rather than the current generation.
- https://huggingface.co/api/models?author=OpenGVLab&search=InternVL&limit=1000 recorded 2026-09-27
155 OpenGVLab InternVL checkpoints with 2,923,773 downloads in the trailing 30 days; InternVL2-1B 709,860 and InternVL2-2B 704,959
Capability
4 high confidenceInternVL3.5 completes tasks in Windows and web environments as well as understanding documents, multi-image input and video. Its online agent scores sit beside UI-TARS-72B in its own report, which places it level with UI-TARS.
- https://arxiv.org/pdf/2508.18265 recorded 2026-09-27
Table 10, GUI grounding and online agentic evaluation: InternVL3.5-241B-A28B ScreenSpot 89.8, ScreenSpot-v2 92.9, OSWorld-G 53.2, WindowsAgentArena 18.0, WebArena-Lite-v2 11.7; UI-TARS-72B 57.1, 17.9, 10.3
- https://huggingface.co/OpenGVLab/InternVL3_5-8B/raw/main/README.md recorded 2026-09-27
"InternVL3.5 supports novel capabilities such as GUI interaction and embodied agency"; multi-image and video inference examples
Verified 2026-09-27