WildGuard
Ai2Ai2's open safety moderation model (7B, fine-tuned from Mistral-7B) that jointly detects harmful prompts, harmful responses, and model refusals across 13 risk subcategories. Trained on the open WildGuardMix corpus and benchmarked above GPT-4 on several moderation tasks.
WildGuard (7B, Mistral-based), Apache-2.0 open weights; trained on the open WildGuardMix dataset. ~265K HF downloads/month, 52 likes (June 2026). Verified live June 2026.
Openness
4 medium confidence- weights
- open(allenai/wildguard on HF)
- data
- open(WildGuardMix corpus public)
- base
- Mistral-7B
- license
- Apache-2.0(OSI)
Open weights under Apache-2.0 with the WildGuardMix training corpus published; from Ai2, which sits at the open end of the spectrum. Scored open_weights (4); a strong open release, just short of the full open_source pipeline tier.
- https://huggingface.co/allenai/wildguard recorded 2026-06-29
Apache-2.0; 7B Mistral fine-tune; prompt/response/refusal detection; open WildGuardMix corpus; ~265K downloads/month
Adoption
4 high confidence~265K HF downloads/month (June 2026); one of the most-downloaded open guardrail models.
- https://huggingface.co/allenai/wildguard recorded 2026-06-29
~265,259 downloads/month, 52 likes (June 2026)
Capability
4 medium confidenceStrong precision and low over-refusal; the joint refusal-detection head is a useful differentiator.
- https://huggingface.co/allenai/wildguard recorded 2026-06-29
prompt/response/refusal detection; 13 subcategories; beats GPT-4 on benchmarks
Unchanged since 2026-06-29 (last edited, not re-checked)