StarCoderData
BigCode ProjectBigCode's filtered code corpus of roughly 250 billion tokens drawn from The Stack v1, covering source files alongside issues and notebooks, used to train StarCoder. It became a standard open code-pretraining component, documented in TinyLlama, IBM Granite Code, ALIA, Apertus, Marin and OLMoE.
Derived from The Stack v1; the map's the-stack-v2 record covers the family's current version. Verified 2026-08-13 via the bigcode/starcoderdata dataset card on Hugging Face.
Openness
3 high confidence- license
- permissive(SPDX per file)
- gated
- terms-acceptance
- dataset_card
- present
Terms-gated like its parent The Stack v1; permissively licensed source files.
- https://huggingface.co/datasets/bigcode/starcoderdata recorded 2026-08-13
gated (auto) access with terms, dataset card
Adoption
3 high confidence28,873 downloads in the trailing 30 days for bigcode/starcoderdata, which bands at 10K-100K, level 3 on the dataset adoption scale.
- https://huggingface.co/api/datasets/bigcode/starcoderdata recorded 2026-08-13
28,873 downloads in the trailing 30 days for bigcode/starcoderdata
Capability
3 high confidenceA proven pretraining component (TinyLlama, Granite Code, ALIA, OLMoE) - the OLMoE mix card lists Starcoder at 101B tokens - but it is drawn from The Stack v1 and functionally superseded by The Stack v2, which is four times its size, and by Stack-Edu.
- https://huggingface.co/datasets/allenai/OLMoE-mix-0924 recorded 2026-08-13
Mix statistics - Starcoder contributes 101B tokens (78.7M docs) to the OLMoE-mix-0924 pretraining mix
- https://arxiv.org/abs/2402.19173 recorded 2026-08-13
Abstract - The Stack v2 is 4x larger than the first StarCoder dataset, and StarCoder2 supersedes the StarCoder lineage
Verified 2026-08-13