AI Potluck
Back to Gap Map Model components / Document conversion & OCR

DeepSeek-OCR

DeepSeek
open weights / Overall score: 4.0(strong)

Vision-language model from DeepSeek that treats a page image as a compressed carrier of text, converting document images into linearised markdown-like output. It was released as a study in optical context compression - how far a page of text can be represented by vision tokens rather than text tokens - and the OCR capability is the demonstration of it.

Openness

3 high confidence
3.0
license
MIT(OSI
weights
open(downloadable from Hugging Face without a gate)
data
closed(the training corpus is neither published nor described)
code
closed(inference code is published

DeepSeek-OCR's code and weights are both released under MIT and download from Hugging Face without a gate, but DeepSeek does not publish or describe the training corpus or the pipeline used to build the model, only the inference code. That makes this an open-weights release rather than a fully open one.

  • https://github.com/deepseek-ai/DeepSeek-OCR recorded 2026-09-16

    Repository page for deepseek-ai/DeepSeek-OCR: MIT, public and unarchived, 23,890 stars, carrying inference code. Establishes the licence, the published code and the absence of a training pipeline.

  • https://huggingface.co/api/models/deepseek-ai/DeepSeek-OCR recorded 2026-09-16

    Hugging Face model API record: license mit, 2,386,735 downloads in the trailing 30 days, weights downloadable without a gate, and no dataset declared. Establishes the open weights and the absent corpus.

Adoption

4 high confidence
4.0

DeepSeek-OCR's adoption is measured by Hugging Face downloads over the trailing month, since Hugging Face is the declared distribution channel for the weights. Docling and MarkItDown, both distributed via PyPI, are measured with higher counts on the same reading.

Capability

not assessed

DeepSeek-OCR converts a page image into text carried in vision tokens, a technique the release calls optical context compression. Neither the repository nor the model card says whether tables or reading order survive that conversion, so no capability score is assigned rather than assumed from silence.

Verified 2026-09-16