AI Potluck
Model components / Training & synthetic datasets

StarCoderData

BigCode Project

StarCoderData is the BigCode project's ~250B-token filtered code corpus drawn from The Stack v1 (permissively licensed GitHub code plus issues and notebooks), used to train StarCoder. It became a standard open code-pretraining component, documented in TinyLlama, IBM Granite Code, ALIA, Apertus, Marin, and OLMoE.

Derived from The Stack v1 (the map's the-stack-v2 entry covers the Stack family's current version; v1 is its direct predecessor).

Openness

3 high confidence
3.0
license
permissive(SPDX per file)
gated
terms-acceptance
dataset_card
present

Terms-gated like its parent The Stack v1; permissively licensed source files.

Adoption

3 high confidence
3.0

20,254 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

3 high confidence
3.0

Proven pretraining component (TinyLlama, Granite Code, ALIA, OLMoE) but drawn from The Stack v1 and functionally superseded by The Stack v2 / Stack-Edu.

Unchanged since 2026-07-04 (last edited, not re-checked)