AI Potluck
Model components / Training & synthetic datasets

The Pile

EleutherAI

EleutherAI's 825 GiB English pretraining corpus from 2020, drawn from 22 sources spanning web text, academic papers, code, books and forums. It set the template for documented open pretraining data and trained Pythia, GPT-NeoX-20B, GPT-J and Cerebras-GPT among others.

A family record covering the original and deduplicated releases. The original mirrors are offline and the data is now served from the deduplicated repository. Verified 2026-08-13 via the EleutherAI/the_pile_deduplicated and EleutherAI/pile dataset cards on Hugging Face.

Openness

3 medium confidence
3.0
access
ungated
license
mixed-per-subset
dataset_card
legacy-repo-only

Ungated download via the deduplicated HF repo; the original EleutherAI/pile repo is a loading script whose mirrors are dead. The card defers licensing to per-subset terms and Books3 has been withdrawn from circulation.

Adoption

3 medium confidence
3.0

The Hugging Face datasets API reports 13,360 downloads in the trailing 30 days for the deduplicated repository EleutherAI/the_pile_deduplicated, ungated and with 116 likes, which bands at 10K-100K, level 3 on the dataset adoption scale. This is a near-boundary reading and should be read as one: 13,360 clears the 10,000 floor of the band by 34%, and the corpus has been sitting close to that floor for months. Confidence is medium for that reason rather than because of anything wrong with the measurement.

Capability

3 high confidence
3.0

The foundational 2020-era template, proven for a wide 2022-23 family (Pythia, GPT-NeoX-20B, GPT-J, Cerebras-GPT), but dated and beaten by modern web corpora in FineWeb ablations.

Verified 2026-08-13