AI Potluck
Model components / Evaluation code

OpenCompass

OpenCompass Community

Comprehensive evaluation platform supporting 100+ datasets across Llama, Mistral, GPT-4, Claude, Qwen, GLM and more, with subjective (arena-style) and objective scoring modes. Picked over lm-eval-harness for Chinese-language and multimodal coverage; backed by Shanghai AI Lab and updated weekly. 7k+ GitHub stars, 162 contributors, 25 commits in last 90 days.

OpenCompass v0.5.2 (Feb 14 2026). 100+ datasets / 70+ with ~400k questions; 20+ HF models + API models (OpenAI/Gemini/Claude). Confirmed live June 2026.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI)
source
public(GitHub)
core-gated
ungated

Apache-2.0, fully public; OSI -> open_source.

Adoption

3 medium confidence
3.0

Widely used eval framework, especially in the China/HF ecosystem; 7.1k GitHub stars and broad dataset+backend coverage. No clean PyPI/download headline located this run; level 3 on reported_traction (framework distribution + star footprint), not on stars alone past the cap.

Capability

4 medium confidence
4.0

Very broad multi-benchmark coverage and backend breadth -- just below lm-eval as the standardization reference; strong but not the single de-facto global standard.

Unchanged since 2026-07-30 (last edited, not re-checked)