AI Potluck
Product / UX / Safety & Guardrails

ShieldGemma

Google

Google DeepMind's content-safety classifier built on Gemma 2, shipped in 2B/9B/27B sizes. It predicts policy violations across four harm categories (sexually explicit, dangerous content, hate, harassment) for both prompts and responses, as a yes/no classifier.

ShieldGemma (2B/9B/27B) on Gemma 2; Gemma license (NOT OSI). 2B variant ~8.5K HF downloads/month, 124 likes (June 2026). Verified live June 2026.

Openness

3 medium confidence
3.0
weights
open(shieldgemma 2B/9B/27B on HF)
data
not-released
license
Gemma(NOT OSI, use-restricted)

Open weights under the Gemma license, which is use-restricted and not OSI-approved (per the openness rubric, Gemma = open_weights, not open_source), so it sits at 3.

Adoption

3 medium confidence
3.0

2B variant ~8.5K HF downloads/month, 124 likes (June 2026); solid adoption across the three sizes.

Capability

3 medium confidence
3.0

Competent four-category classifier with a size ladder; narrower hazard taxonomy than Llama Guard / Granite Guardian.

Unchanged since 2026-06-29 (last edited, not re-checked)