AI Potluck
Model components / Fine-tuned / chat models

StarCoder2

BigCode Project

Scored on StarCoder2-15B-Instruct-v0.1 (the instruct SKU; 16B, self-aligned, released 31 Oct 2024, arXiv:2410.24198). NOTE: base StarCoder2-15B (Feb 2024) is a pretrained model belonging in base_pretrained; only the instruct variant fits finetuned_chat. Both verified live on HF June 2026.

Scored on StarCoder2-15B-Instruct-v0.1 (the instruct SKU; 16B, self-aligned, released 31 Oct 2024, arXiv:2410.24198). NOTE: base StarCoder2-15B (Feb 2024) is a pretrained model belonging in base_pretrained; only the instruct variant fits finetuned_chat. Both verified live on HF June 2026.

Openness

4 high confidence
4.0
weights
open
data
open(self-oss-instruct-sc2-exec-filter-50k + 7 documented intermediate datasets)
code
open(starcoder2-self-align pipeline on GitHub)
base-data
open(The Stack v2)
license
BigCode-OpenRAIL-M(open but RAIL behavioral-use restrictions, not OSI)

The most open artifact in this category by evidence: weights, the full training corpus (The Stack v2) and the self-align pipeline are all published. Held to 4 rather than 5 because BigCode-OpenRAIL-M is not OSI-approved - its behavioral-use clauses discriminate against fields of endeavor, which the OSD forbids - and rule 5 of the ladder reserves the top score for an OSI license. Class corrected from open_source to open_weights on 2026-07-29: the score was right but 4/open_source is a pair this ladder cannot emit, and the mismatch is what flagged it. Resolved under issue #117 by tiering RAIL as a conduct restriction, which is what keeps a fully-open pipeline off the 2/restricted floor it would otherwise have hit on the license alone.

Adoption

1 medium confidence
1.0

Instruct SKU ~1.1k downloads/last-month; base StarCoder2-15B ~9.4k/last-month. Research-significant (fully-open self-alignment recipe) but low production reach; placed at 1 on the instruct SKU's measured volume.

Capability

3 high confidence
3.0

Strong HumanEval for a fully-open 16B in 2024 (best fully-self-aligned code model at release), but HumanEval is saturated and the model is below 2026 frontier coders on harder agentic benchmarks (SWE-bench). Mid-tier.

Unchanged since 2026-07-29 (last edited, not re-checked)