AI Potluck
Back to Gap Map Model components / Multimodal models

North Micro Vision

Cohere
open weights / Overall score: 2.3

North Micro Vision is Cohere's 2.4B-parameter vision-language model, part of its North family of purpose-built models. It reads interleaved text and images at native resolution and handles question answering, captioning, grounding, OCR, charts and documents in several languages. The vision encoder is trained from SigLIP 2 and the language model is Cohere's own.

Openness

3 high confidence
3.0
weights
open(model.safetensors on the Hub, ungated)
data
described(publicly available datasets and an in-house multilingual document corpus, named in the release post
code
partial(Transformers inference, a NeMo AutoModel fine-tuning recipe made with NVIDIA and community Axolotl support
license
Apache-2.0(OSI

North Micro Vision's weights are released under the Apache 2.0 license and are not gated. Cohere trained it partly on its own document corpus and publishes neither that data nor its training code.

Adoption

3 high confidence
3.0

Hugging Face downloads over the trailing 30 days of North Micro Vision Instruct, the only North vision checkpoint.

Capability

2 high confidence
2.0

North Micro Vision reads documents and charts well for its size and takes several images at once. Its card claims no video understanding and rules out tool use, so it sits level with Moondream.

Verified 2026-09-27