AI Potluck
Model components / Dataset Processing Tools

olmOCR

Allen Institute for AI

Converts PDFs and images to clean markdown/text using a 7B vision-language model, handling equations, tables, and handwriting. Emits Dolma-format documents; used in the Dolma 3 PDF-extraction path.

Ai2 olmOCR; repo active (v0.4.27, 2026-03-12), Apache-2.0 confirmed live 2026-07-21. High-profile document-extraction tool (19k+ GitHub stars); feeds the Dolma 3 pipeline.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI)
source
public(allenai/olmocr)
governance
Allen Institute for AI
core-gated
ungated

Fully OSI-licensed (Apache-2.0), full source public.

Adoption

3 high confidence
3.0

~22k PyPI downloads/month and 19k+ GitHub stars; a leading open document-extraction tool, used in the Dolma 3 data path.

Capability

4 high confidence
4.0

Benchmark-anchored: leads open document-extraction tools on olmOCR-Bench but sits mid-pack on the neutral OmniDocBench, so near-SOTA rather than SOTA. Snapshot; extraction leaderboards move.

Unchanged since 2026-07-30 (last edited, not re-checked)