AI Potluck
Product / UX / Safety & Guardrails

HarmBench

Center for AI Safety

A standardized evaluation framework for automated red-teaming from the Center for AI Safety. It runs a common pipeline across many attack methods and target models so red-team attacks and defenses can be compared rigorously and co-developed.

HarmBench, from the Center for AI Safety. A standardized red-teaming evaluation that shipped covering 18 attack methods against 33 target models. Verified 2026-08-13 via GitHub.

Openness

5 high confidence
5.0
license
MIT(OSI)
source
public
self-host
yes
service
none
core-gated
ungated

Fully open source: MIT-licensed standardized red-teaming evaluation framework, self-hostable.

  • https://github.com/centerforaisafety/HarmBench recorded 2026-08-13

    Repo metadata records license spdxId - MIT. The whole framework is in the public repo and the README is a Quick Start for running it yourself; there is no enterprise or commercial edition, no paid tier and no license key anywhere in the page, so nothing is withheld from the published source.

Adoption

2 low confidence
2.0

1,029 stargazers on the HarmBench repo, which bands at 1K-10K stars, level 2, and stars are the only signal available: HarmBench ships no package on PyPI or npm, so there is no download route to band on instead. A stars proxy never rises above 3, because a star is not a use. It is widely cited in research as the standardized red-teaming benchmark.

Capability

4 medium confidence
4.0

The reference standardized harness for comparing red-team attacks and defenses; research-grade breadth.

  • https://github.com/centerforaisafety/HarmBench recorded 2026-08-13

    README describes 'a fast, scalable, and open-source framework for evaluating automated red teaming methods and LLM attacks/defenses' and records the initial release as covering 33 evaluated LLMs and 18 red teaming methods.

Verified 2026-08-13