PaddleOCR
PaddlePaddleOCR and document-parsing toolkit from Baidu's PaddlePaddle project, covering text detection and recognition in more than a hundred languages plus table, formula and seal recognition, and a document-parsing pipeline that emits Markdown and JSON. PaddleOCR-VL is its vision-language parsing model.
Openness
5 high confidence- license
- Apache-2.0(OSI)
- source
- public(PaddlePaddle/PaddleOCR is the toolkit you run)
- core features withheld
- no — a 2,968-entry recursive tree carries no path matching enterprise, ee, commercial or proprietary
PaddleOCR is Apache-2.0 throughout, confirmed from the LICENSE body. The full repository tree carries no enterprise, ee or commercial path, and the toolkit is published as part of Baidu's PaddlePaddle open-source project rather than sold as a separate edition.
- https://api.github.com/repos/PaddlePaddle/PaddleOCR/git/trees/main?recursive=1 recorded 2026-09-16
Full untruncated recursive tree of the default branch, 2,968 entries. One root LICENSE. No path matches enterprise, ee, commercial or proprietary. A tree lists paths; it is cited for that and for nothing about the vendor's offerings.
- https://github.com/PaddlePaddle/PaddleOCR recorded 2026-09-16
The PaddleOCR repository page. PaddleOCR is published as part of Baidu's PaddlePaddle open-source project with no commercial edition offered; cited for the licence, the public source and the absence of a paid tier.
Adoption
4 high confidenceAdoption is measured as PyPI downloads of the paddleocr package over the trailing month.
- https://pypistats.org/api/packages/paddleocr/recent recorded 2026-09-16
Trailing-window download counts for the declared PyPI artifact: 1,495,891 downloads in the trailing month.
Capability
5 medium confidencePaddleOCR's PP-ChatOCRv4 pipeline returns the fields a user names from a document, on top of a parsing pipeline that emits Markdown and JSON with tables, formulas and seals recovered, which matches Docling's template-driven extraction. Its support for more than 100 languages is not what places it here, since language breadth measures how many documents it can process, not how much of one survives.
- https://github.com/PaddlePaddle/PaddleOCR recorded 2026-09-16
Repository page and README describing multilingual text detection and recognition, table, equation and seal recognition, and the document-parsing pipeline.
- https://raw.githubusercontent.com/PaddlePaddle/PaddleOCR/main/docs/version3.x/pipeline_usage/PP-ChatOCRv4.en.md recorded 2026-09-29
The PP-ChatOCRv4-doc tutorial describes a pipeline built "to address complex document information extraction challenges such as layout analysis, rare characters, multi-page PDFs, tables, and seal text recognition", imports it with "from paddleocr import PPChatOCRv4Doc", passes the fields wanted as "key_list=[" and prints the extracted field as "{'chat_res': {'驾驶室准乘人数': '2'}}".
- https://raw.githubusercontent.com/PaddlePaddle/PaddleOCR/main/README.md recorded 2026-09-29
The raw README states "PaddleOCR converts PDF documents and images into structured, LLM-ready data (JSON/Markdown)" and that "additional dependencies for document parsing and information extraction can be installed as needed".
Verified 2026-09-16