AI Potluck
Model components / Dataset Processing Tools

datatrove

Hugging Face

Platform-agnostic library of composable pipeline blocks for large-scale text data processing (reading, filtering, deduplication, tokenization) that runs unchanged across local, Slurm, and cloud backends. The reference pretraining-data pipeline behind the FineWeb family.

Built FineWeb, FineWeb-Edu, FineWeb-2 and FineMath; the de-facto open pretraining-data processing pipeline. Verified 2026-08-13 via the GitHub README and the LICENSE body.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI)
source
public(huggingface/datatrove)
governance
Hugging Face
managed-tier
none
core-gated
ungated

Fully OSI-licensed (Apache-2.0), full source public, no proprietary core or managed gate.

Adoption

2 high confidence
2.0

65,009 downloads in the trailing 30 days for the declared `datatrove` package, which bands at 10K-100K, level 2 on the software and model adoption scale.

Capability

5 high confidence
5.0

Benchmark-anchored on the corpus datatrove builds, as a proxy for the tool itself: FineWeb-Edu leads the public pretraining corpora on the FineWeb ablation suite. The arXiv abstract supports that ranking claim; the per-task ablation figures recorded here are plotted in the paper rather than printed on the abstract page.

  • https://arxiv.org/abs/2406.17557 recorded 2026-08-13

    FineWeb abstract - FineWeb produces better-performing LLMs than other open pretraining datasets, and FineWeb-Edu lifts knowledge and reasoning benchmarks such as MMLU and ARC dramatically. The per-task ablation figures behind the aggregate are plots in the paper rather than text on this page.

Verified 2026-08-13