AI Potluck
Back to Gap Map Model components / Document conversion & OCR

Chandra

Datalab
source available / Overall score: 2.2

Document OCR and layout model from Datalab that converts PDFs and images into Markdown, HTML and JSON while preserving structure - reading order, tables, forms, handwriting and mathematical notation. It is the successor model to the same team's Marker pipeline and is served both as open weights and as a hosted API.

The code and the weights carry different licences; see the openness note.

Openness

2 high confidence
2.0
license
Apache-2.0(OSI
source
public(datalab-to/chandra carries the inference code and the weights are published)

Chandra's code is Apache-2.0 and fully public, but its model weights carry a modified AI Pubs Open RAIL-M license: free for research, personal use, and startups under $2M in funding or revenue, and barred to anyone whose product competes with Datalab's. That revenue threshold and anti-compete clause, on the weights rather than the code, keep this at source-available rather than fully open. The license on the weights allows commercial use only below the threshold, so its tier is competition_restricted.

  • https://api.github.com/repos/datalab-to/chandra/git/trees/master?recursive=1 recorded 2026-09-16

    Full untruncated recursive tree of the default branch, 70 entries. One root LICENSE. No path matches enterprise, ee, commercial or proprietary. A tree lists paths; it is cited for that and for nothing about the vendor's offerings.

  • https://github.com/datalab-to/chandra recorded 2026-09-16

    Repository page for datalab-to/chandra: public and unarchived, 12,267 stars, label Apache-2.0. Establishes that the source is public.

  • https://huggingface.co/datalab-to/chandra-ocr-2/raw/main/LICENSE recorded 2026-09-28

    The weights LICENSE is headed "AI PUBS OPEN RAIL-M LICENSE (MODIFIED) / Version 0.1, March 2, 2023 (Modified)". Attachment A bars use "for any purpose if You ... generated more than two million US Dollars ($2,000,000) in gross revenue in the prior year, except where Your Use is limited to personal use or research purposes", the same for USD 2M raised, and use by anyone who "provides or otherwise makes available any product or service that competes with any product or service offered by or made available by Licensor".

  • https://raw.githubusercontent.com/datalab-to/chandra/master/chandra/settings.py recorded 2026-09-28

    Chandra's settings default the weights to MODEL_CHECKPOINT: str = "datalab-to/chandra-ocr-2".

  • https://raw.githubusercontent.com/datalab-to/chandra/master/README.md recorded 2026-09-16

    Chandra's README: two badges, 'Code License Apache-2.0' and 'Model License OpenRAIL-M', and under Commercial usage - 'This code is Apache 2.0, and our model weights use a modified OpenRAIL-M license (free for research, personal use, and startups under $2M funding/revenue, cannot be used competitively with our API). To remove the OpenRAIL license requirements, or for broader commercial licensing, visit our pricing page'. Establishes the two-licence split and the terms on the weights.

Adoption

1 high confidence
1.0

Chandra's adoption is read from PyPI downloads of chandra-ocr over the trailing month, the only usage signal recorded for the package. The reading lands just under the threshold for the next reach bracket up.

Capability

4 medium confidence
4.0

Chandra recovers reading order, tables, forms, handwriting and mathematical notation into Markdown, HTML and JSON, but stops short of Docling, whose extractor also returns the fields a user defines in a template. It tops the olmOCR-Bench leaderboard, ahead of olmOCR itself.

  • https://github.com/datalab-to/chandra recorded 2026-09-16

    Repository page and README describing PDF and image conversion to Markdown, HTML and JSON with layout, tables, forms, handwriting and mathematical notation.

  • https://raw.githubusercontent.com/datalab-to/chandra/master/README.md recorded 2026-09-29

    The raw README describes "a state of the art OCR model that converts images and PDFs into structured HTML/Markdown/JSON while preserving layout information" and lists "Tops external olmocr benchmark and significant improvement in internal multilingual benchmarks", "Reconstructs forms accurately, including checkboxes" and "Strong performance with tables, math, and complex layouts".

Verified 2026-09-16