AI Potluck
Model components / Dataset Processing Tools

Dolma toolkit

Allen Institute for AI

High-performance toolkit for curating LM pretraining data: taggers, quality and content filtering, and Rust Bloom-filter deduplication. The software that built the Dolma corpus and the OLMo/OLMoE mixes.

Ai2 Dolma toolkit (allenai/dolma), distinct from the Dolma dataset artifact. Built the Dolma corpus and the OLMo/OLMoE pretraining mixes. Verified 2026-08-13 via the GitHub README and the LICENSE body.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI)
source
public(allenai/dolma)
governance
Allen Institute for AI
core-gated
ungated

Fully OSI-licensed (Apache-2.0), full source public.

Adoption

1 high confidence
1.0

4,390 downloads in the trailing 30 days for dolma, read from the pypistats API, which bands at <10K, level 1 on the software and model adoption scale. It is the curation toolkit behind the Dolma corpus and the OLMo family, so its importance runs ahead of its raw download volume.

Capability

3 high confidence
3.0

Benchmark-anchored on the corpus the toolkit builds, as a proxy for the tool: capable and fully open, but the released Dolma corpus trails the datatrove and NeMo Curator frontier on DataComp-LM.

  • https://arxiv.org/html/2406.11794v3 recorded 2026-08-13

    DataComp-LM Table 2 in the full text - Dolma-V1 scores 35.0 Core and 18.4 Extended at the 7B-1x scale, below RedPajama 35.3 and RefinedWeb 36.9

Verified 2026-08-13