AI Potluck
Model components / Training & synthetic datasets

The Stack v2

BigCode Project

The Stack v2 is a large source-code dataset of 3.28B unique files (67.53TB uncompressed) from 104.2M GitHub repositories spanning 658 programming languages, sourced via Software Heritage. It was created by the BigCode project and used to train the StarCoder2 models. Files carry permissive licenses, but access requires accepting terms of use tied to Software Heritage principles.

Verified live 2026-06-22 via primary sources. Requires acceptance of terms and sharing contact info before download, so access is gated.

Openness

3 high confidence
3.0
license
permissive(SPDX)
gated
terms-acceptance+contact-info
dataset_card
yes

Requires acceptance of terms and sharing contact info before download, so access is gated.

Adoption

3 high confidence
3.0

28,469 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

4 high confidence
4.0

Trained StarCoder2 (2024) and is the code base of SmolLM3 (2025) plus the source for Stack-Edu; as the raw corpus it is beaten by its own edu-filtered slice in ablations.

Unchanged since 2026-07-04 (last edited, not re-checked)