olmOCR
Allen Institute for AIConverts PDFs and images to clean markdown/text using a 7B vision-language model, handling equations, tables, and handwriting. Emits Dolma-format documents; used in the Dolma 3 PDF-extraction path.
Ai2 olmOCR; repo active (v0.4.27, 2026-03-12), Apache-2.0 confirmed live 2026-07-21. High-profile document-extraction tool (19k+ GitHub stars); feeds the Dolma 3 pipeline.
Openness
5 high confidence- license
- Apache-2.0(OSI)
- source
- public(allenai/olmocr)
- governance
- Allen Institute for AI
- core-gated
- ungated
Fully OSI-licensed (Apache-2.0), full source public.
- https://github.com/allenai/olmocr/blob/main/LICENSE recorded 2026-07-21
Apache License Version 2.0 text
Adoption
3 high confidence~22k PyPI downloads/month and 19k+ GitHub stars; a leading open document-extraction tool, used in the Dolma 3 data path.
- https://pypistats.org/packages/olmocr recorded 2026-07-21
Downloads last month: 21,753
Capability
4 high confidenceBenchmark-anchored: leads open document-extraction tools on olmOCR-Bench but sits mid-pack on the neutral OmniDocBench, so near-SOTA rather than SOTA. Snapshot; extraction leaderboards move.
- https://github.com/allenai/olmocr recorded 2026-07-21
olmOCR 2 (v0.4.0): 82.4 on olmOCR-Bench, ahead of Marker/MinerU/Mistral OCR/GPT-4o
- https://github.com/opendatalab/OmniDocBench recorded 2026-07-21
OmniDocBench: olmOCR ~85.7, mid-pack behind MinerU2.5-Pro and Gemini-3 Pro
Unchanged since 2026-07-30 (last edited, not re-checked)