The Stack v2
BigCode ProjectThe Stack v2 is a large source-code dataset of 3.28B unique files (67.53TB uncompressed) from 104.2M GitHub repositories spanning 658 programming languages, sourced via Software Heritage. It was created by the BigCode project and used to train the StarCoder2 models. Files carry permissive licenses, but access requires accepting terms of use tied to Software Heritage principles.
Verified live 2026-06-22 via primary sources. Requires acceptance of terms and sharing contact info before download, so access is gated.
Openness
3 high confidence- license
- permissive(SPDX)
- gated
- terms-acceptance+contact-info
- dataset_card
- yes
Requires acceptance of terms and sharing contact info before download, so access is gated.
- https://huggingface.co/datasets/bigcode/the-stack-v2 recorded 2026-06-22
gated access requiring terms acceptance and contact-info sharing, permissive licenses, dataset card
Adoption
3 high confidence28,469 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/bigcode/the-stack-v2 recorded 2026-07-04
downloads field = 28,469 (30-day)
Capability
4 high confidenceTrained StarCoder2 (2024) and is the code base of SmolLM3 (2025) plus the source for Stack-Edu; as the raw corpus it is beaten by its own edu-filtered slice in ablations.
- https://arxiv.org/abs/2402.19173 recorded 2026-07-04
StarCoder2 trained on The Stack v2 (900B+ tokens, 619 languages)
- https://huggingface.co/blog/smollm3 recorded 2026-07-04
SmolLM3 code data is The Stack v2
Unchanged since 2026-07-04 (last edited, not re-checked)