Docling
IBMOpen-source SDK and CLI that parses PDF, DOCX, PPTX, XLSX, HTML, EPUB, images, and audio into unified structured output (Markdown, lossless JSON, HTML, DocTags) for gen-AI and RAG pipelines. Unlike web-URL-to-Markdown services (Firecrawl, Jina Reader), it does deep local document understanding of rich binary docs - layout/reading-order analysis, table-structure recognition, formula/code extraction, OCR, and VLM integration - running locally as a library.
MIT; originated at IBM Research, inducted into the Linux Foundation AI & Data Foundation (Docling Project) on 2025-04-29 under community governance. ~62k stars (exceptionally high).
Openness
5 high confidence- license
- MIT(OSI)
- source
- public(github.com/docling-project/docling)
- self-host
- primary
- core-gated
- ungated
MIT, local-first document parsing, fully self-hostable with integrated open-weight models.
- https://github.com/docling-project/docling recorded 2026-06-18
MIT, ~61.8k stars, multi-format document parser
Adoption
3 medium confidence~62k stars (exceptionally high) but PyPI download volume not surfaced - stars_fallback (capped at 3; true usage almost certainly higher).
- https://github.com/docling-project/docling recorded 2026-06-18
~61.8k stars
Capability
5 high confidenceDeepest local document understanding among RAG ingestion tools; web-URL parsers lack rich binary-doc depth.
- https://github.com/docling-project/docling recorded 2026-06-18
feature set: multi-format, layout, tables, OCR, VLM
Unchanged since 2026-07-30 (last edited, not re-checked)