AI Potluck
Model components / Evaluation code

HELM

Stanford CRFM

Holistic Evaluation of Language Models, Python framework for reproducible, transparent evaluation of foundation models including LLMs and multimodal models, organized around scenarios, adapters, and metrics. Picked for academic-grade reproducibility and breadth (accuracy, robustness, fairness, bias, toxicity). 2.8k stars, 128 contributors, 151 commits in last 90 days; Stanford has announced HELM will enter maintenance mode June 2026.

stanford-crfm/helm, v0.5.16 (released 30 Apr 2026), Apache-2.0, ~2.8k stars. Standardized holistic evaluation framework + public leaderboards: HELM Capabilities, HELM Safety, VHELM (vision-language), HEIM (text-to-image), MedHELM, audio-LM, enterprise benchmarks; supports MMLU-Pro, GPQA, IFEval, WildBench. Confirmed live June 2026.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI)
source
public(framework + standardized datasets + web UI + leaderboard code)
governance
Stanford CRFM(academic)
core-gated
ungated

Fully OSI (Apache-2.0); framework, standardized data formats, web UI, and leaderboards all public.

Adoption

3 medium confidence
3.0

Headline signal is influence/standard-setting, not install volume: HELM is a widely-cited standardized public leaderboard from Stanford CRFM spanning multiple modalities (capabilities/safety/vision/medical), referenced across the research community. PyPI distribution (crfm-helm) is only ~3,142 downloads/month; most usage is via the hosted leaderboard and academic citation rather than pip, so download volume understates reach. Level 3 reflects meaningful research/standard adoption short of mass production deployment.

Capability

5 high confidence
5.0

Frontier-defining on coverage and standardization, the holistic multi-modality leaderboard standard. One of the 1-2 category leaders, scored 5. (Weaker than Inspect AI on agentic-eval support, but broader on benchmark/modality coverage.)

  • https://github.com/stanford-crfm/helm recorded 2026-06-04

    standardized datasets, unified model access, metrics beyond accuracy, web UI, multi-modal leaderboards (MMLU-Pro/GPQA/IFEval/WildBench)

Unchanged since 2026-07-30 (last edited, not re-checked)