BLIP
SalesforceBLIP is Salesforce's family of vision-language models for captioning, visual question answering and image-text matching. BLIP-2 connects a frozen image encoder to a frozen language model through a lightweight querying transformer, and InstructBLIP adds instruction tuning. xGen-MM, also called BLIP-3, continues the line with Phi-3-based models that read interleaved images and text; BLIP-3-Video is a separate video model that Salesforce built on it.
The LAVIS repository, which holds the BLIP, BLIP-2 and xGen-MM training code on separate branches, is archived and names no successor. BLIP-3-Video and Salesforce's image-generation models BLIP3-o and BLIP3o-NEXT are not part of this entry.
Openness
3 high confidence- weights
- open(safetensors for the xGen-MM v1.5 base and instruct checkpoints on the Hub, ungated)
- data
- partial(Salesforce publishes BLIP3-KALE, BLIP3-OCR-200M and BLIP3-GROUNDING-50M, three of the datasets it built for training
- code
- open(the xgen-mm branch of LAVIS, "the fine-tuning code that's used for producing our instruct models")
- license
- Apache-2.0(OSI
- video-model
- restricted(xGen-MM-Vid (BLIP-3-Video), a separate video model built on the v1.5 architecture, ships under CC-BY-NC-4.0)
xGen-MM v1.5, the BLIP-3 release, is Apache 2.0, and Salesforce publishes the code that fine-tuned its instruct models along with three of the large datasets it built for training. The instruction-tuning mixture itself is not released, so the model cannot be rebuilt exactly.
- https://arxiv.org/abs/2408.08872 recorded 2026-09-27
"Our training code, models, and all datasets used in this work, including the three largescale datasets we create and the preprocessed ones, will be open-sourced"
- https://cdn.jsdelivr.net/gh/salesforce/LAVIS@xgen-mm/README.md recorded 2026-09-27
"This codebase provides the fine-tuning code that's used for producing our instrcut models"; "We implement xGen-MM in the OpenFlamingo codebase"
- https://huggingface.co/api/datasets?author=Salesforce&search=blip3&limit=1000&expand[]=downloads&expand[]=createdAt&expand[]=cardData&expand[]=gated&expand[]=lastModified recorded 2026-09-27
blip3-kale, blip3-ocr-200m and blip3-grounding-50m published by Salesforce, each tagged apache-2.0
- https://huggingface.co/api/models?author=Salesforce&search=xgen-mm&limit=1000&expand[]=downloads&expand[]=createdAt&expand[]=cardData&expand[]=gated&expand[]=lastModified recorded 2026-09-27
Eight xGen-MM checkpoints: the v1.5 base and instruct models tagged apache-2.0, the v1 instruct model and the two xgen-mm-vid models cc-by-nc-4.0; the newest is xgen-mm-vid, 2025-01-15
- https://huggingface.co/api/models/Salesforce/xgen-mm-phi3-mini-instruct-interleave-r-v1.5?expand[]=downloads&expand[]=cardData&expand[]=gated&expand[]=createdAt&expand[]=lastModified&expand[]=siblings recorded 2026-09-27
Hub metadata for xgen-mm-phi3-mini-instruct-interleave-r-v1.5: gated false, safetensors shards, card license apache-2.0
- https://huggingface.co/Salesforce/xgen-mm-phi3-mini-instruct-interleave-r-v1.5/raw/main/README.md recorded 2026-09-27
"In the v1.5 (08/2024) release"; interleave is "our main instruct model"; "This series advances upon the successful designs of the BLIP series"; "Our code and weights are released under the Apache 2.0 license"; "This release is for research purposes only"
- https://huggingface.co/Salesforce/xgen-mm-vid-phi3-mini-r-v1.5-128tokens-8frames/raw/main/README.md recorded 2026-09-27
xGen-MM-Vid (BLIP-3-Video) "specifically designed to understand videos"; "Our code and weights are released under the CC by-NC 4.0 license"
- https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/salesforce%2FLAVIS recorded 2026-09-27
salesforce/LAVIS: archived true, fork false, license bsd-3-clause
Adoption
4 high confidenceHugging Face downloads over the trailing 30 days, summed across Salesforce's BLIP, BLIP-2, InstructBLIP and xGen-MM checkpoints. Most of it is the original BLIP captioning and question-answering models, which remain common pipeline components.
- https://huggingface.co/api/models?author=Salesforce&limit=1000&expand[]=downloads&expand[]=createdAt&expand[]=cardData&expand[]=gated&expand[]=lastModified recorded 2026-09-27
BLIP checkpoints 2,846,161, BLIP-2 813,835, InstructBLIP 65,012 and xGen-MM 1,800 downloads in the trailing 30 days, 3,726,808 in all; blip-image-captioning-base 1,625,727
Capability
2 medium confidencexGen-MM, the current BLIP model, reads interleaved sequences of several images and the text in them, beyond BLIP-2's single image. It reports no document or chart benchmark and takes no video, so it sits level with Moondream; Salesforce's separate BLIP-3-Video model is not counted here.
- https://arxiv.org/abs/2301.12597 recorded 2026-09-27
BLIP-2 "outperforms Flamingo80B by 8.7% on zero-shot VQAv2"
- https://arxiv.org/abs/2408.08872 recorded 2026-09-27
"with the ability to comprehend interleaved image-text inputs"
- https://huggingface.co/Salesforce/xgen-mm-phi3-mini-instruct-interleave-r-v1.5/raw/main/README.md recorded 2026-09-27
Multi-image benchmarks for xGen-MM-inst.-interleave (4B): BLINK 49.7, QBench-2 75.1, Mantis-eval 56.7; single-image MMMU (val) 41.1, TextVQA 71.0
Verified 2026-09-27