Tesseract OCR
Tesseract OCRLong-standing open-source OCR engine, originally developed at HP and maintained since 2006 with Google's support. It recognises text in over a hundred languages with an LSTM-based engine, and outputs plain text, hOCR, ALTO, TSV and PDF. It remains the default OCR dependency for a large amount of other software rather than a document-conversion product in its own right.
No package channel this map routes - it is distributed as source and distribution packages, and the popular pytesseract PyPI package is a third-party wrapper rather than the product - so adoption is read from stars.
Openness
5 high confidence- license
- Apache-2.0(OSI)
- source
- public(tesseract-ocr/tesseract is the engine)
- core features withheld
- no — a 780-entry recursive tree carries no path matching enterprise, ee, commercial or proprietary, and no vendor sells a Tesseract edition
Apache-2.0 with no vendor and no paid edition. The 780-entry recursive tree carries no enterprise, ee or commercial path.
- https://api.github.com/repos/tesseract-ocr/tesseract/git/trees/main?recursive=1 recorded 2026-09-16
Full untruncated recursive tree of the default branch, 780 entries. One root LICENSE. No path matches enterprise, ee, commercial or proprietary. A tree lists paths; it is cited for that and for nothing about the vendor's offerings.
- https://github.com/tesseract-ocr/tesseract recorded 2026-09-16
The Tesseract repository page. Tesseract is maintained as a community open-source project with no vendor and no commercial edition; cited for the licence, the public source and the absence of a paid tier.
Adoption
3 low confidenceTesseract is distributed as source and through operating-system packages rather than through a registry this map can measure, leaving GitHub stars as the only comparable adoption signal. That almost certainly understates the use of a library embedded in a very large amount of other software.
- https://github.com/tesseract-ocr/tesseract recorded 2026-09-16
Repository page showing the star count for tesseract-ocr/tesseract.
Capability
3 medium confidenceTesseract's LSTM engine recognizes characters and their positions in plain text, hOCR, ALTO, TSV or PDF, but recovers no reading order, tables or document structure of its own. It remains the OCR engine many other document tools build their own recognition on, even though it leaves structure recovery to whatever wraps it.
- https://github.com/tesseract-ocr/tesseract recorded 2026-09-16
Repository page and README describing the LSTM OCR engine, its language coverage and its output formats.
Verified 2026-09-16