ShieldGemma
GoogleGoogle DeepMind's content-safety classifier built on Gemma 2, shipped in 2B/9B/27B sizes. It predicts policy violations across four harm categories (sexually explicit, dangerous content, hate, harassment) for both prompts and responses, as a yes/no classifier.
ShieldGemma (2B/9B/27B) on Gemma 2; Gemma license (NOT OSI). 2B variant ~8.5K HF downloads/month, 124 likes (June 2026). Verified live June 2026.
Openness
3 medium confidence- weights
- open(shieldgemma 2B/9B/27B on HF)
- data
- not-released
- license
- Gemma(NOT OSI, use-restricted)
Open weights under the Gemma license, which is use-restricted and not OSI-approved (per the openness rubric, Gemma = open_weights, not open_source), so it sits at 3.
- https://huggingface.co/google/shieldgemma-2b recorded 2026-06-29
Gemma license (not OSI); open weights; four harm categories; ~8,555 downloads/month, 124 likes
Adoption
3 medium confidence2B variant ~8.5K HF downloads/month, 124 likes (June 2026); solid adoption across the three sizes.
- https://huggingface.co/google/shieldgemma-2b recorded 2026-06-29
~8,555 downloads/month, 124 likes (June 2026)
Capability
3 medium confidenceCompetent four-category classifier with a size ladder; narrower hazard taxonomy than Llama Guard / Granite Guardian.
- https://huggingface.co/google/shieldgemma-2b recorded 2026-06-29
four harm categories; 2B/9B/27B sizes
Unchanged since 2026-06-29 (last edited, not re-checked)