AI Potluck
Product / UX / Safety & Guardrails

HarmBench

Center for AI Safety

A standardized evaluation framework for automated red-teaming from the Center for AI Safety. It runs a common pipeline across many attack methods and target models so red-team attacks and defenses can be compared rigorously and co-developed.

HarmBench, MIT, Center for AI Safety; ~1K GitHub stars (June 2026). Standardized red-teaming eval across 18 attacks and 33 target models. Verified live June 2026.

Openness

5 high confidence
5.0
license
MIT(OSI)
source
public
self-host
yes
service
none
core-gated
ungated

Fully open source: MIT-licensed standardized red-teaming evaluation framework, self-hostable.

Adoption

2 low confidence
2.0

~1K GitHub stars (June 2026); stars-only proxy. Widely cited as the standardized red-teaming benchmark in research.

Capability

4 medium confidence
4.0

The reference standardized harness for comparing red-team attacks and defenses; research-grade breadth.

Unchanged since 2026-08-01 (last edited, not re-checked)