AI Potluck
Back to Gap Map Model components / Document conversion & OCR

PaddleOCR

PaddlePaddle
open source / Overall score: 4.4(strong)

OCR and document-parsing toolkit from Baidu's PaddlePaddle project, covering text detection and recognition in more than a hundred languages plus table, formula and seal recognition, and a document-parsing pipeline that emits Markdown and JSON. PaddleOCR-VL is its vision-language parsing model.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI)
source
public(PaddlePaddle/PaddleOCR is the toolkit you run)
core features withheld
no — a 2,968-entry recursive tree carries no path matching enterprise, ee, commercial or proprietary

PaddleOCR is Apache-2.0 throughout, confirmed from the LICENSE body. The full repository tree carries no enterprise, ee or commercial path, and the toolkit is published as part of Baidu's PaddlePaddle open-source project rather than sold as a separate edition.

  • https://api.github.com/repos/PaddlePaddle/PaddleOCR/git/trees/main?recursive=1 recorded 2026-09-16

    Full untruncated recursive tree of the default branch, 2,968 entries. One root LICENSE. No path matches enterprise, ee, commercial or proprietary. A tree lists paths; it is cited for that and for nothing about the vendor's offerings.

  • https://github.com/PaddlePaddle/PaddleOCR recorded 2026-09-16

    The PaddleOCR repository page. PaddleOCR is published as part of Baidu's PaddlePaddle open-source project with no commercial edition offered; cited for the licence, the public source and the absence of a paid tier.

Adoption

4 high confidence
4.0

Adoption is measured as PyPI downloads of the paddleocr package over the trailing month.

Capability

5 medium confidence
5.0

PaddleOCR's PP-ChatOCRv4 pipeline returns the fields a user names from a document, on top of a parsing pipeline that emits Markdown and JSON with tables, formulas and seals recovered, which matches Docling's template-driven extraction. Its support for more than 100 languages is not what places it here, since language breadth measures how many documents it can process, not how much of one survives.

Verified 2026-09-16