HELM
Stanford CRFMHolistic Evaluation of Language Models, Python framework for reproducible, transparent evaluation of foundation models including LLMs and multimodal models, organized around scenarios, adapters, and metrics. Picked for academic-grade reproducibility and breadth (accuracy, robustness, fairness, bias, toxicity). 2.8k stars, 128 contributors, 151 commits in last 90 days; Stanford has announced HELM will enter maintenance mode June 2026.
stanford-crfm/helm, v0.5.16 (released 30 Apr 2026), Apache-2.0, ~2.8k stars. Standardized holistic evaluation framework + public leaderboards: HELM Capabilities, HELM Safety, VHELM (vision-language), HEIM (text-to-image), MedHELM, audio-LM, enterprise benchmarks; supports MMLU-Pro, GPQA, IFEval, WildBench. Confirmed live June 2026.
Openness
5 high confidence- license
- Apache-2.0(OSI)
- source
- public(framework + standardized datasets + web UI + leaderboard code)
- governance
- Stanford CRFM(academic)
- core-gated
- ungated
Fully OSI (Apache-2.0); framework, standardized data formats, web UI, and leaderboards all public.
- https://github.com/stanford-crfm/helm recorded 2026-06-04
Apache-2.0 license; holistic eval framework with standardized datasets, web UI, leaderboards; v0.5.16; 2.8k stars
Adoption
3 medium confidenceHeadline signal is influence/standard-setting, not install volume: HELM is a widely-cited standardized public leaderboard from Stanford CRFM spanning multiple modalities (capabilities/safety/vision/medical), referenced across the research community. PyPI distribution (crfm-helm) is only ~3,142 downloads/month; most usage is via the hosted leaderboard and academic citation rather than pip, so download volume understates reach. Level 3 reflects meaningful research/standard adoption short of mass production deployment.
- https://github.com/stanford-crfm/helm recorded 2026-06-04
multiple public leaderboards (HELM Capabilities/Safety/VHELM/HEIM/MedHELM); Stanford CRFM standardized benchmark
- https://pypistats.org/packages/crfm-helm recorded 2026-06-04
3,142 PyPI downloads last month (low because usage is leaderboard/citation-driven, not pip)
Capability
5 high confidenceFrontier-defining on coverage and standardization, the holistic multi-modality leaderboard standard. One of the 1-2 category leaders, scored 5. (Weaker than Inspect AI on agentic-eval support, but broader on benchmark/modality coverage.)
- https://github.com/stanford-crfm/helm recorded 2026-06-04
standardized datasets, unified model access, metrics beyond accuracy, web UI, multi-modal leaderboards (MMLU-Pro/GPQA/IFEval/WildBench)
Unchanged since 2026-07-30 (last edited, not re-checked)