AI Potluck
Back to Gap Map Model components / Document conversion & OCR

Surya

Datalab
source available / Overall score: 3.4

OCR, layout analysis and reading-order engine from Datalab covering more than ninety languages. It detects text lines, classifies layout blocks, orders them for reading, and recognises tables - and it is the engine underneath the same team's Marker document pipeline rather than a competing end-to-end converter.

It is the engine Marker runs on, which is why the two are scored as separate products rather than one.

Openness

2 medium confidence
2.0
license
Apache-2.0(OSI
covers the code)+AI-Pubs-Open-RAIL-M-Modified(the model weights
datalab-to/surya-ocr-2 and surya_layout2, the default recognition and layout checkpoints
source
public(datalab-to/surya carries the engine and the weights are published)
ocr-error-model
cc-by-nc-sa-4.0(the Hub copy of Datalab's OCR-error detector, which the settings load from an s3:// path whose license is not stated

Apache-2.0 covers Surya's code, while the recognition and layout weights it loads by default carry a modified AI Pubs Open RAIL-M license: free for research, personal use and startups under $5M in funding or revenue, and barred to anyone whose product competes with Datalab's. A separate error-detection model is tagged non-commercial on the Hub and would land on the same score. The license on the weights allows commercial use only below the threshold, so its tier is competition_restricted and the product reads as source-available.

  • https://api.github.com/repos/datalab-to/surya/git/trees/master?recursive=1 recorded 2026-09-16

    Full untruncated recursive tree of the default branch, 205 entries. One root LICENSE. No path matches enterprise, ee, commercial or proprietary. A tree lists paths; it is cited for that and for nothing about the vendor's offerings.

  • https://github.com/datalab-to/surya recorded 2026-09-16

    Repository page for datalab-to/surya: public and unarchived, 21,393 stars, label Apache-2.0. Establishes that the source is public.

  • https://huggingface.co/api/models/datalab-to/ocr_error_detection recorded 2026-09-28

    Hub metadata for datalab-to/ocr_error_detection: "license":"cc-by-nc-sa-4.0", "gated":false.

  • https://huggingface.co/datalab-to/surya-ocr-2/raw/main/LICENSE recorded 2026-09-28

    The weights LICENSE is headed "AI PUBS OPEN RAIL-M LICENSE (MODIFIED) / Version 0.1, March 2, 2023 (Modified)". Attachment A bars use "for any purpose if You ... generated more than five million US Dollars ($5,000,000) in gross revenue in the prior year, except where Your Use is limited to personal use or research purposes", the same for USD 5M raised, and use by anyone who "provides or otherwise makes available any product or service that competes with any product or service offered by or made available by Licensor".

  • https://raw.githubusercontent.com/datalab-to/surya/master/README.md recorded 2026-09-16

    Surya's README carries the same two Datalab badges as Chandra - 'Code License Apache-2.0' and 'Model License OpenRAIL-M' linking to the Datalab pricing page. Establishes that the weights are not under the code's licence.

  • https://raw.githubusercontent.com/datalab-to/surya/master/surya/settings.py recorded 2026-09-28

    Surya's settings default SURYA_MODEL_CHECKPOINT: str = "datalab-to/surya-ocr-2" and FAST_LAYOUT_MODEL_CHECKPOINT: str = "hf://datalab-to/surya_layout2", with OCR_ERROR_MODEL_CHECKPOINT: str = "s3://ocr_error_detection/2025_02_18".

Adoption

3 high confidence
3.0

Adoption is measured as PyPI downloads of the surya-ocr package over the trailing month.

Capability

4 medium confidence
4.0

Surya performs text line detection, layout block classification, reading-order ordering and table recognition. Being an engine that other tools call, rather than a finished converter, does not reduce how much of a document it recovers, which puts it a step above Tesseract.

  • https://github.com/datalab-to/surya recorded 2026-09-16

    Repository page and README describing OCR, line detection, layout analysis, reading order and table recognition across more than ninety languages.

Verified 2026-09-16