StarCoderData
BigCode ProjectStarCoderData is the BigCode project's ~250B-token filtered code corpus drawn from The Stack v1 (permissively licensed GitHub code plus issues and notebooks), used to train StarCoder. It became a standard open code-pretraining component, documented in TinyLlama, IBM Granite Code, ALIA, Apertus, Marin, and OLMoE.
Derived from The Stack v1 (the map's the-stack-v2 entry covers the Stack family's current version; v1 is its direct predecessor).
Openness
3 high confidence- license
- permissive(SPDX per file)
- gated
- terms-acceptance
- dataset_card
- present
Terms-gated like its parent The Stack v1; permissively licensed source files.
- https://huggingface.co/datasets/bigcode/starcoderdata recorded 2026-07-04
gated (auto) access with terms, dataset card
Adoption
3 high confidence20,254 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/bigcode/starcoderdata recorded 2026-07-04
downloads field = 20,254 (30-day)
Capability
3 high confidenceProven pretraining component (TinyLlama, Granite Code, ALIA, OLMoE) but drawn from The Stack v1 and functionally superseded by The Stack v2 / Stack-Edu.
- https://huggingface.co/datasets/allenai/OLMoE-mix-0924 recorded 2026-07-04
StarCoderData is a documented component of TinyLlama, Granite Code, ALIA, Apertus, Marin, OLMoE
- https://arxiv.org/abs/2402.19173 recorded 2026-07-04
The Stack v2/StarCoder2 is the successor to the Stack v1 lineage
Unchanged since 2026-07-04 (last edited, not re-checked)