The Stack v2
BigCode ProjectSource-code dataset of 3.28 billion unique files, 67.53TB uncompressed, drawn from 104.2 million GitHub repositories across 658 programming languages and sourced through Software Heritage. It was created by the BigCode project and used to train the StarCoder2 models.
Verified 2026-08-13 via the bigcode/the-stack-v2 dataset card on Hugging Face.
Openness
3 high confidence- license
- permissive(SPDX)
- gated
- terms-acceptance+contact-info
- dataset_card
- yes
Requires acceptance of terms and sharing contact info before download, so access is gated.
- https://huggingface.co/datasets/bigcode/the-stack-v2 recorded 2026-08-13
gated access requiring terms acceptance and contact-info sharing, permissive licenses, dataset card
Adoption
3 high confidence13,060 downloads in the trailing 30 days for bigcode/the-stack-v2, which bands at 10K-100K, level 3 on the dataset adoption scale.
- https://huggingface.co/api/datasets/bigcode/the-stack-v2 recorded 2026-08-13
13,060 downloads in the trailing 30 days for bigcode/the-stack-v2
Capability
4 high confidenceTrained StarCoder2 (2024) and is the code base of SmolLM3 (2025), as well as the source for Stack-Edu. The corpus spans 619 programming languages and is four times the size of the first StarCoder dataset, but as raw code it is beaten in ablation by its own edu-filtered slice, which is why it sits a band below Stack-Edu.
- https://arxiv.org/abs/2402.19173 recorded 2026-08-13
Abstract - The Stack v2 spans 619 programming languages and is 4x the first StarCoder dataset; StarCoder2 3B/7B/15B train on 3.3-4.3T tokens of it
- https://huggingface.co/blog/smollm3 recorded 2026-08-13
SmolLM3 code data is The Stack v2
Verified 2026-08-13