AI Potluck
Back to Gap Map Model components / Document conversion & OCR

dots.ocr

Studio Dots AI
open weights / Overall score: 3.4

Multilingual document-parsing model that performs layout detection and content recognition in a single vision-language model rather than a pipeline of separate stages, switching task by prompt. It emits layout blocks with their reading order and the text inside them.

The GitHub API reports the repository under studio-dots-ai; it is commonly cited under a rednote-hilab path that does not resolve, and the LICENSE body carries a rednote-hilab copyright line, so the project moved organizations. Typed `model` as a judgement rather than by a rule: the release is one vision-language model, and the repository is inference tooling for it, shipping an installable package, setup.py, that the model card installs with pip install -e .

Openness

3 high confidence
3.0
license
MIT(OSI
weights
open(published on Hugging Face and downloadable without a gate)
data
closed(the training corpus is neither published nor described)
code
partial(inference code and a loader are published

MIT covers both the code and the weights, which download freely from Hugging Face. But the repository publishes inference code only, not the training pipeline, and the training corpus is neither published nor described, so this reads as an open-weights release rather than a fully open model.

  • https://github.com/studio-dots-ai/dots.ocr recorded 2026-09-16

    Repository page for studio-dots-ai/dots.ocr: MIT, public and unarchived, 9,117 stars, carrying inference code with the weights published on Hugging Face. Establishes the licence, the published code and the absence of a training pipeline.

Adoption

3 high confidence
3.0

Adoption is measured as Hugging Face downloads of the dots.ocr checkpoint over the trailing month, the only usage-volume channel this record declares.

Capability

4 medium confidence
4.0

dots.ocr's model card documents layout detection, reading order and table recognition, all emitted as JSON. Its capability matches Chandra's, the other single-model parser here.

  • https://github.com/studio-dots-ai/dots.ocr recorded 2026-09-16

    Repository page and README describing unified layout detection and content recognition in one vision-language model with prompt-based task switching, and table recognition emitted as JSON.

Verified 2026-09-16