AI Potluck
Back to Gap Map Model components / Multimodal models

PaliGemma

Google
open weights / Overall score: 2.3

PaliGemma is Google's vision-language model built from the SigLIP vision encoder and Gemma language models, designed to be fine-tuned for specific tasks. PaliGemma 2 comes in 3B, 10B and 28B sizes at 224 to 896 pixel resolution and outputs captions, answers, text read from images, bounding boxes and segmentation codes. Google releases pretrained, task-transfer and mixed-task checkpoints.

Openness

3 high confidence
3.0
weights
open(safetensors on the Hub behind a click-through license acknowledgement)
data
described(WebLI, CC3M-35L, OpenImages and WIT are named for pretraining
code
open(the PaliGemma fine-tuning trainer and transfer configs in google-research/big_vision)
license
Gemma-License(the Gemma Terms of Use, which list PaliGemma and PaliGemma 2 and incorporate the Prohibited Use Policy)

PaliGemma is released under the Gemma Terms of Use, which carry Google's prohibited-use policy and require passing its restrictions on to anyone downstream. Google publishes the fine-tuning code but not the pretraining data.

Adoption

3 high confidence
3.0

Hugging Face downloads over the trailing 30 days, summed across Google's PaliGemma and PaliGemma 2 checkpoints. Most of it is the first PaliGemma, and downloads from Kaggle and Vertex AI are not counted.

Capability

2 high confidence
2.0

PaliGemma reads text, charts and tables and returns boxes and segmentation masks. Its general mix checkpoints take one image; video and multi-image results come only from checkpoints fine-tuned for a single task, so it sits level with Moondream.

  • https://arxiv.org/abs/2412.03555 recorded 2026-09-26

    Abstract: table structure recognition, molecular structure recognition, music score recognition, long fine-grained captioning and radiography reports

  • https://huggingface.co/google/paligemma2-3b-mix-224 recorded 2026-09-26

    "a caption of the image, an answer to a question, a list of object bounding box coordinates, or segmentation codewords"; transfer table DocVQA 76.6, ChartQA 66.4, AI2D 84.4 at 448px 10B

Verified 2026-09-26