PaliGemma
GooglePaliGemma is Google's vision-language model built from the SigLIP vision encoder and Gemma language models, designed to be fine-tuned for specific tasks. PaliGemma 2 comes in 3B, 10B and 28B sizes at 224 to 896 pixel resolution and outputs captions, answers, text read from images, bounding boxes and segmentation codes. Google releases pretrained, task-transfer and mixed-task checkpoints.
Openness
3 high confidence- weights
- open(safetensors on the Hub behind a click-through license acknowledgement)
- data
- described(WebLI, CC3M-35L, OpenImages and WIT are named for pretraining
- code
- open(the PaliGemma fine-tuning trainer and transfer configs in google-research/big_vision)
- license
- Gemma-License(the Gemma Terms of Use, which list PaliGemma and PaliGemma 2 and incorporate the Prohibited Use Policy)
PaliGemma is released under the Gemma Terms of Use, which carry Google's prohibited-use policy and require passing its restrictions on to anyone downstream. Google publishes the fine-tuning code but not the pretraining data.
- https://ai.google.dev/gemma/terms recorded 2026-09-26
"Gemma Terms of Use"; the Appendix lists PaliGemma and PaliGemma 2; section 3.2 incorporates the Gemma Prohibited Use Policy
- https://huggingface.co/api/models/google/paligemma2-3b-mix-224?expand[]=downloads&expand[]=cardData&expand[]=gated&expand[]=createdAt&expand[]=lastModified&expand[]=siblings recorded 2026-09-26
Hub metadata for paligemma2-3b-mix-224: gated manual (license acknowledgement), two safetensors shards, card license gemma
- https://huggingface.co/google/paligemma2-3b-mix-224 recorded 2026-09-26
The card names WebLI, CC3M-35L, VQ2A-CC3M-35L, OpenImages and WIT as pretraining data; "fine-tuned on a mixture of academic tasks"; no dataset link
- https://raw.githubusercontent.com/google-research/big_vision/main/big_vision/configs/proj/paligemma/README.md recorded 2026-09-26
"How to run PaliGemma fine-tuning" with big_vision.trainers.proj.paligemma.train and transfer configs
- https://raw.githubusercontent.com/google-research/big_vision/main/big_vision/trainers/proj/paligemma/train.py recorded 2026-09-26
The PaliGemma trainer in big_vision
Adoption
3 high confidenceHugging Face downloads over the trailing 30 days, summed across Google's PaliGemma and PaliGemma 2 checkpoints. Most of it is the first PaliGemma, and downloads from Kaggle and Vertex AI are not counted.
- https://huggingface.co/api/models?author=google&search=paligemma&limit=1000 recorded 2026-09-26
166 Google PaliGemma checkpoints with 651,853 downloads in the trailing 30 days; PaliGemma 2's 43 checkpoints account for 61,830
Capability
2 high confidencePaliGemma reads text, charts and tables and returns boxes and segmentation masks. Its general mix checkpoints take one image; video and multi-image results come only from checkpoints fine-tuned for a single task, so it sits level with Moondream.
- https://arxiv.org/abs/2412.03555 recorded 2026-09-26
Abstract: table structure recognition, molecular structure recognition, music score recognition, long fine-grained captioning and radiography reports
- https://huggingface.co/google/paligemma2-3b-mix-224 recorded 2026-09-26
"a caption of the image, an answer to a question, a list of object bounding box coordinates, or segmentation codewords"; transfer table DocVQA 76.6, ChartQA 66.4, AI2D 84.4 at 448px 10B
Verified 2026-09-26