AI Potluck
Back to Gap Map Model components / Document conversion & OCR

Tesseract OCR

Tesseract OCR
open source / Overall score: 3.0

Long-standing open-source OCR engine, originally developed at HP and maintained since 2006 with Google's support. It recognises text in over a hundred languages with an LSTM-based engine, and outputs plain text, hOCR, ALTO, TSV and PDF. It remains the default OCR dependency for a large amount of other software rather than a document-conversion product in its own right.

No package channel this map routes - it is distributed as source and distribution packages, and the popular pytesseract PyPI package is a third-party wrapper rather than the product - so adoption is read from stars.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI)
source
public(tesseract-ocr/tesseract is the engine)
core features withheld
no — a 780-entry recursive tree carries no path matching enterprise, ee, commercial or proprietary, and no vendor sells a Tesseract edition

Apache-2.0 with no vendor and no paid edition. The 780-entry recursive tree carries no enterprise, ee or commercial path.

Adoption

3 low confidence
3.0

Tesseract is distributed as source and through operating-system packages rather than through a registry this map can measure, leaving GitHub stars as the only comparable adoption signal. That almost certainly understates the use of a library embedded in a very large amount of other software.

Capability

3 medium confidence
3.0

Tesseract's LSTM engine recognizes characters and their positions in plain text, hOCR, ALTO, TSV or PDF, but recovers no reading order, tables or document structure of its own. It remains the OCR engine many other document tools build their own recognition on, even though it leaves structure recovery to whatever wraps it.

Verified 2026-09-16