HarmBench
Center for AI SafetyA standardized evaluation framework for automated red-teaming from the Center for AI Safety. It runs a common pipeline across many attack methods and target models so red-team attacks and defenses can be compared rigorously and co-developed.
HarmBench, from the Center for AI Safety. A standardized red-teaming evaluation that shipped covering 18 attack methods against 33 target models. Verified 2026-08-13 via GitHub.
Openness
5 high confidence- license
- MIT(OSI)
- source
- public
- self-host
- yes
- service
- none
- core-gated
- ungated
Fully open source: MIT-licensed standardized red-teaming evaluation framework, self-hostable.
- https://github.com/centerforaisafety/HarmBench recorded 2026-08-13
Repo metadata records license spdxId - MIT. The whole framework is in the public repo and the README is a Quick Start for running it yourself; there is no enterprise or commercial edition, no paid tier and no license key anywhere in the page, so nothing is withheld from the published source.
Adoption
2 low confidence1,029 stargazers on the HarmBench repo, which bands at 1K-10K stars, level 2, and stars are the only signal available: HarmBench ships no package on PyPI or npm, so there is no download route to band on instead. A stars proxy never rises above 3, because a star is not a use. It is widely cited in research as the standardized red-teaming benchmark.
- https://github.com/centerforaisafety/HarmBench recorded 2026-08-13
stargazerCount 1029 for centerforaisafety/HarmBench; no download, install or customer figure appears on the page.
Capability
4 medium confidenceThe reference standardized harness for comparing red-team attacks and defenses; research-grade breadth.
- https://github.com/centerforaisafety/HarmBench recorded 2026-08-13
README describes 'a fast, scalable, and open-source framework for evaluating automated red teaming methods and LLM attacks/defenses' and records the initial release as covering 33 evaluated LLMs and 18 red teaming methods.
Verified 2026-08-13