AI Potluck
Model components / Dataset Processing Tools

olmOCR

Allen Institute for AI

Converts PDFs and images to clean markdown/text using a 7B vision-language model, handling equations, tables, and handwriting. Emits Dolma-format documents; used in the Dolma 3 PDF-extraction path.

Feeds the Dolma 3 PDF-extraction path and ships its own olmOCR-Bench suite. Verified 2026-08-13 via the GitHub README and the LICENSE body.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI)
source
public(allenai/olmocr)
governance
Allen Institute for AI
core-gated
ungated

Fully OSI-licensed (Apache-2.0), full source public.

Adoption

2 high confidence
2.0

22,925 PyPI downloads in the trailing 30 days for olmocr, which bands at 10K-100K, level 2 on the software and model adoption scale.

Capability

4 high confidence
4.0

Benchmark-anchored: olmOCR scores near the top of olmOCR-Bench, the benchmark the project publishes itself, ahead of Marker, MinerU and Mistral OCR but now behind Chandra OCR and Infinity-Parser; on the third-party OmniDocBench it sits mid-pack. Near-SOTA rather than SOTA, and a snapshot at that, because extraction leaderboards move.

Verified 2026-08-13