ShieldGemma
GoogleGoogle DeepMind's content-safety classifier built on Gemma 2, shipped in 2B/9B/27B sizes. It predicts policy violations across four harm categories (sexually explicit, dangerous content, hate, harassment) for both prompts and responses, as a yes/no classifier.
ShieldGemma (2B/9B/27B) on Gemma 2, an English-only yes/no classifier over four harm categories. The weights are behind an access request rather than open to download, and the card publishes no training data. Verified 2026-08-13 via the HF model card.
Openness
3 medium confidence- weights
- open(shieldgemma 2B/9B/27B on HF)
- data
- not-released
- license
- Gemma(NOT OSI, use-restricted)
Open weights under the Gemma license, which carries use restrictions and is not OSI-approved. Gemma terms count as open weights rather than open source throughout this map, so ShieldGemma sits at 3.
- https://huggingface.co/google/shieldgemma-2b recorded 2026-08-13
Repo shows License - gemma, gated manual, and two safetensors shards, so the weights are downloadable after accepting the Gemma terms. The card describes the models as 'available in English with open weights' and declares no training datasets.
- https://huggingface.co/api/models/google/shieldgemma-2b recorded 2026-08-13
tags include license:gemma; gated - manual; siblings list model-00001-of-00002.safetensors and model-00002-of-00002.safetensors; cardData carries no datasets key.
Adoption
1 medium confidence4,761 downloads in the trailing 30 days for google/shieldgemma-2b, which bands at under 10K, level 1. A guardrail model with modest standalone pull-through; most use comes through the Gemma tooling rather than through direct download.
- https://huggingface.co/api/models/google/shieldgemma-2b recorded 2026-08-12
4,761 downloads in the trailing 30 days for google/shieldgemma-2b
Capability
3 medium confidenceA competent four-category classifier with a size ladder, but a narrower hazard taxonomy than Llama Guard or Granite Guardian - four categories against Llama Guard's fourteen, which puts it one rung below.
- https://huggingface.co/google/shieldgemma-2b recorded 2026-08-13
Card describes 'a series of safety content moderation models built upon Gemma 2 that target four harm categories (sexually explicit, dangerous content, hate, and harassment)', in 2B, 9B and 27B sizes, English only.
Verified 2026-08-12