peS2o
Allen Institute for AIAi2's corpus of roughly 40 million cleaned academic papers derived from S2ORC on Semantic Scholar. It is a standard open scientific-text component, documented in the pretraining mixes for OLMoE, Marin and NVIDIA's Nemotron 3.
It contributes about 57 billion tokens to the OLMoE pretraining mix. Verified 2026-08-13 via the allenai/peS2o dataset card on Hugging Face.
Openness
5 high confidence- license
- ODC-BY
- gated
- false
- dataset_card
- present
ODC-BY, ungated, full dataset card.
- https://huggingface.co/datasets/allenai/peS2o recorded 2026-08-13
ODC-BY license, ungated, dataset card
Adoption
3 high confidence11,867 downloads in the trailing 30 days for allenai/peS2o, which bands at 10K-100K, level 3 on the dataset adoption scale.
- https://huggingface.co/api/datasets/allenai/peS2o recorded 2026-08-13
11,867 downloads in the trailing 30 days for allenai/peS2o
Capability
4 high confidenceThe standard open science corpus with documented multi-model reuse (OLMoE 57B tokens, Marin, Nemotron 3), but no ablation leadership, and OLMo 3 replaced it with a new academic-PDF dataset. The OLMoE mix card lists peS2o at 57.2B tokens and the Nemotron 3 Nano card names it among its pretraining sources; neither claims an ablation win for it.
- https://huggingface.co/datasets/allenai/OLMoE-mix-0924 recorded 2026-08-13
Mix statistics - peS2o (Dolma) contributes 57.2B tokens (38.8M docs) to the OLMoE-mix-0924 pretraining mix
- https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 recorded 2026-08-13
Nemotron 3 Nano lists peS2o among pretraining sources
Verified 2026-08-13