AI Potluck
Back to Gap Map Model components / Training & synthetic datasets

Wikipedia

Wikimedia Foundation
open / Overall score: 4.0(strong)

Cleaned article text from the Wikipedia dumps in more than 300 languages, one subset per language, published as Parquet in the wikimedia namespace on Hugging Face. The current release is a single dump from November 2023. Wikipedia text is a small but widely reused component of pretraining mixtures.

Openness

5 high confidence
5.0
license
CC-BY-SA-4.0(the grant dumps.wikimedia.org/legal.html now states for text)+cc-by-sa-3.0(the grant the card and Hub metadata still state, covering older text
gated
false
dataset_card
present
format
parquet

Wikipedia text is under Creative Commons Attribution-ShareAlike. The Foundation's legal page for the dumps now states 4.0, while the card and Hub metadata still name 3.0; both allow redistribution and commercial reuse with attribution and share-alike. Most text is also under the GFDL, but some is available only under Creative Commons, so the CC grants are what cover the release. The download is ungated, and the card documents structure and fields, though its curation, bias and limitations sections are empty.

  • https://dumps.wikimedia.org/legal.html recorded 2026-09-26

    "all original textual content is licensed under the GNU Free Documentation License (GFDL) and the Creative Commons Attribution-Share-Alike 4.0 License. Some text may be available only under the Creative Commons license"

  • https://huggingface.co/api/datasets/wikimedia/wikipedia recorded 2026-09-26

    Hub API metadata for the dataset, license tags cc-by-sa-3.0 and gfdl, gated false, published by the wikimedia organization

  • https://huggingface.co/datasets/wikimedia/wikipedia/raw/main/README.md recorded 2026-09-26

    dataset card with summary, languages, structure, fields, splits and creation sections; "All original textual content is licensed under the GNU Free Documentation License (GFDL) and the Creative Commons Attribution-Share-Alike 3.0 License. Some text may be available only under the Creative Commons license"

Adoption

4 medium confidence
4.0

Adoption is measured by Hugging Face downloads of wikimedia/wikipedia, the cleaned text published in the wikimedia namespace on the Hub, over the trailing month. It excludes the other ways people obtain Wikipedia, such as the raw dumps on dumps.wikimedia.org.

Capability

4 high confidence
4.0

Wikipedia text is a documented part of major pretraining mixes. LLaMA sampled it at 4.5 percent, about 2.45 epochs over its 1.4-trillion-token run, and OLMoE's mix carries it through Dolma. That evidence is about Wikipedia text, not this particular Hugging Face distribution of it. No ablation shows it beating a current web corpus, so the score comes from documented reuse rather than a measured gain.

Verified 2026-09-26