AI Potluck
Back to Gap Map Model components / Document conversion & OCR

LightOnOCR

LightOn
open weights / Overall score: 3.0

Compact vision-language OCR model from LightOn, released at around one billion parameters and trained end-to-end to turn page images into markdown-formatted text. Its argument is throughput: it targets the accuracy of far larger document models at a fraction of the inference cost.

Distributed as Hugging Face weights rather than through a package registry.

Openness

3 high confidence
3.0
license
Apache-2.0(OSI
weights
open(downloadable from Hugging Face without a gate)
data
closed(the training corpus is not published)
code
closed(no training pipeline is published

LightOnOCR's weights are released under Apache-2.0 and download from Hugging Face without a gate, but LightOn does not publish the training corpus or the pipeline used to produce the model, only the weights themselves. That makes this an open-weights release rather than a fully open one.

  • https://huggingface.co/api/models/lightonai/LightOnOCR-2-1B recorded 2026-09-16

    Hugging Face model API record for lightonai/LightOnOCR-2-1B: license apache-2.0, 129,301 downloads in the trailing 30 days, 807 likes, weights downloadable without a gate, and no training dataset declared. Establishes the licence, the open weights and the absent recipe.

Adoption

3 high confidence
3.0

LightOnOCR's adoption is measured by Hugging Face downloads of the model over the trailing month, the declared distribution channel for the weights.

Capability

not assessed

LightOnOCR's Hugging Face record establishes its size, license and download count, but says nothing about whether tables or reading order survive conversion to markdown. Parameter count and inference speed are not a substitute for that, so no capability score is assigned.

Verified 2026-09-16