AI Potluck
Model components / Training & synthetic datasets

The Stack v2

BigCode Project

Source-code dataset of 3.28 billion unique files, 67.53TB uncompressed, drawn from 104.2 million GitHub repositories across 658 programming languages and sourced through Software Heritage. It was created by the BigCode project and used to train the StarCoder2 models.

Verified 2026-08-13 via the bigcode/the-stack-v2 dataset card on Hugging Face.

Openness

3 high confidence
3.0
license
permissive(SPDX)
gated
terms-acceptance+contact-info
dataset_card
yes

Requires acceptance of terms and sharing contact info before download, so access is gated.

Adoption

3 high confidence
3.0

13,060 downloads in the trailing 30 days for bigcode/the-stack-v2, which bands at 10K-100K, level 3 on the dataset adoption scale.

Capability

4 high confidence
4.0

Trained StarCoder2 (2024) and is the code base of SmolLM3 (2025), as well as the source for Stack-Edu. The corpus spans 619 programming languages and is four times the size of the first StarCoder dataset, but as raw code it is beaten in ablation by its own edu-filtered slice, which is why it sits a band below Stack-Edu.

Verified 2026-08-13