LightOnOCR
LightOnCompact vision-language OCR model from LightOn, released at around one billion parameters and trained end-to-end to turn page images into markdown-formatted text. Its argument is throughput: it targets the accuracy of far larger document models at a fraction of the inference cost.
Distributed as Hugging Face weights rather than through a package registry.
Openness
3 high confidence- license
- Apache-2.0(OSI
- weights
- open(downloadable from Hugging Face without a gate)
- data
- closed(the training corpus is not published)
- code
- closed(no training pipeline is published
LightOnOCR's weights are released under Apache-2.0 and download from Hugging Face without a gate, but LightOn does not publish the training corpus or the pipeline used to produce the model, only the weights themselves. That makes this an open-weights release rather than a fully open one.
- https://huggingface.co/api/models/lightonai/LightOnOCR-2-1B recorded 2026-09-16
Hugging Face model API record for lightonai/LightOnOCR-2-1B: license apache-2.0, 129,301 downloads in the trailing 30 days, 807 likes, weights downloadable without a gate, and no training dataset declared. Establishes the licence, the open weights and the absent recipe.
Adoption
3 high confidenceLightOnOCR's adoption is measured by Hugging Face downloads of the model over the trailing month, the declared distribution channel for the weights.
- https://huggingface.co/api/models/lightonai/LightOnOCR-2-1B recorded 2026-09-16
Hugging Face model API record: 129,301 downloads in the trailing 30 days.
Capability
not assessedLightOnOCR's Hugging Face record establishes its size, license and download count, but says nothing about whether tables or reading order survive conversion to markdown. Parameter count and inference speed are not a substitute for that, so no capability score is assigned.
- https://huggingface.co/api/models/lightonai/LightOnOCR-2-1B recorded 2026-09-16
The model's Hugging Face record: what is published, its size class and the declared licence, and no statement about what document structure the output preserves.
Verified 2026-09-16