AI Potluck
Model components / Training & synthetic datasets

The Pile

EleutherAI

The Pile is EleutherAI's 825 GiB, 22-source English pretraining corpus (2020), spanning web text, academic papers, code, books, and forums. It set the template for documented open pretraining data and trained Pythia, GPT-NeoX-20B, GPT-J, and Cerebras-GPT among others. The original mirrors are offline; data is now served through the deduplicated Hugging Face repo, and component licenses vary by subset.

Covers the Pile family (original + deduplicated). Verified live 2026-07-04; canonical hosting moved to the deduplicated repo after the original mirrors went offline.

Openness

3 medium confidence
3.0
access
ungated
license
mixed-per-subset
dataset_card
legacy-repo-only

Ungated download via the deduplicated HF repo; the original EleutherAI/pile repo is a loading script whose mirrors are dead. The card defers licensing to per-subset terms and Books3 has been withdrawn from circulation.

Adoption

2 high confidence
2.0

9,980 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

3 high confidence
3.0

The foundational 2020-era template, proven for a wide 2022-23 family (Pythia, GPT-NeoX-20B, GPT-J, Cerebras-GPT), but dated and beaten by modern web corpora in FineWeb ablations.

Unchanged since 2026-07-04 (last edited, not re-checked)