AI Potluck
Product / UX / Safety & Guardrails

WildGuard

Ai2

Ai2's open safety moderation model (7B, fine-tuned from Mistral-7B) that jointly detects harmful prompts, harmful responses, and model refusals across 13 risk subcategories. Trained on the open WildGuardMix corpus and benchmarked above GPT-4 on several moderation tasks.

WildGuard (7B, Mistral-based), Apache-2.0 open weights; trained on the open WildGuardMix dataset. ~265K HF downloads/month, 52 likes (June 2026). Verified live June 2026.

Openness

4 medium confidence
4.0
weights
open(allenai/wildguard on HF)
data
open(WildGuardMix corpus public)
base
Mistral-7B
license
Apache-2.0(OSI)

Open weights under Apache-2.0 with the WildGuardMix training corpus published; from Ai2, which sits at the open end of the spectrum. Scored open_weights (4); a strong open release, just short of the full open_source pipeline tier.

Adoption

4 high confidence
4.0

~265K HF downloads/month (June 2026); one of the most-downloaded open guardrail models.

Capability

4 medium confidence
4.0

Strong precision and low over-refusal; the joint refusal-detection head is a useful differentiator.

Unchanged since 2026-06-29 (last edited, not re-checked)