AI Potluck
Model components / Training & synthetic datasets

StarCoderData

BigCode Project

BigCode's filtered code corpus of roughly 250 billion tokens drawn from The Stack v1, covering source files alongside issues and notebooks, used to train StarCoder. It became a standard open code-pretraining component, documented in TinyLlama, IBM Granite Code, ALIA, Apertus, Marin and OLMoE.

Derived from The Stack v1; the map's the-stack-v2 record covers the family's current version. Verified 2026-08-13 via the bigcode/starcoderdata dataset card on Hugging Face.

Openness

3 high confidence
3.0
license
permissive(SPDX per file)
gated
terms-acceptance
dataset_card
present

Terms-gated like its parent The Stack v1; permissively licensed source files.

Adoption

3 high confidence
3.0

28,873 downloads in the trailing 30 days for bigcode/starcoderdata, which bands at 10K-100K, level 3 on the dataset adoption scale.

Capability

3 high confidence
3.0

A proven pretraining component (TinyLlama, Granite Code, ALIA, OLMoE) - the OLMoE mix card lists Starcoder at 101B tokens - but it is drawn from The Stack v1 and functionally superseded by The Stack v2, which is four times its size, and by Stack-Edu.

Verified 2026-08-13