AI Potluck
Model components / Training & synthetic datasets

peS2o

Allen Institute for AI

Ai2's corpus of roughly 40 million cleaned academic papers derived from S2ORC on Semantic Scholar. It is a standard open scientific-text component, documented in the pretraining mixes for OLMoE, Marin and NVIDIA's Nemotron 3.

It contributes about 57 billion tokens to the OLMoE pretraining mix. Verified 2026-08-13 via the allenai/peS2o dataset card on Hugging Face.

Openness

5 high confidence
5.0
license
ODC-BY
gated
false
dataset_card
present

ODC-BY, ungated, full dataset card.

Adoption

3 high confidence
3.0

11,867 downloads in the trailing 30 days for allenai/peS2o, which bands at 10K-100K, level 3 on the dataset adoption scale.

Capability

4 high confidence
4.0

The standard open science corpus with documented multi-model reuse (OLMoE 57B tokens, Marin, Nemotron 3), but no ablation leadership, and OLMo 3 replaced it with a new academic-PDF dataset. The OLMoE mix card lists peS2o at 57.2B tokens and the Nemotron 3 Nano card names it among its pretraining sources; neither claims an ablation win for it.

Verified 2026-08-13