datatrove
Hugging FacePlatform-agnostic library of composable pipeline blocks for large-scale text data processing (reading, filtering, deduplication, tokenization) that runs unchanged across local, Slurm, and cloud backends. The reference pretraining-data pipeline behind the FineWeb family.
Hugging Face datatrove; repo active (v0.9.0, 2026-03-04), Apache-2.0 confirmed live 2026-07-21. Built FineWeb, FineWeb-Edu, FineWeb-2, and FineMath; the de-facto open pretraining-data processing pipeline.
Openness
5 high confidence- license
- Apache-2.0(OSI)
- source
- public(huggingface/datatrove)
- governance
- Hugging Face
- managed-tier
- none
- core-gated
- ungated
Fully OSI-licensed (Apache-2.0), full source public, no proprietary core or managed gate.
- https://github.com/huggingface/datatrove/blob/main/LICENSE recorded 2026-07-21
Apache License Version 2.0 text
Adoption
3 high confidence~55k PyPI downloads/month; the reference pipeline for the FineWeb dataset family and widely reused across open pretraining-data efforts.
- https://pypistats.org/packages/datatrove recorded 2026-07-21
Downloads last month: 55,132
Capability
5 high confidenceBenchmark-anchored on the output corpus (a proxy for the tool): FineWeb-Edu leads public pretraining corpora on the FineWeb ablation suite.
- https://arxiv.org/abs/2406.17557 recorded 2026-07-21
FineWeb ablation: FineWeb-Edu tops public datasets on an 8-task aggregate (identical 1.82B model, equal tokens)
Unchanged since 2026-07-30 (last edited, not re-checked)