North Micro Vision
CohereNorth Micro Vision is Cohere's 2.4B-parameter vision-language model, part of its North family of purpose-built models. It reads interleaved text and images at native resolution and handles question answering, captioning, grounding, OCR, charts and documents in several languages. The vision encoder is trained from SigLIP 2 and the language model is Cohere's own.
Openness
3 high confidence- weights
- open(model.safetensors on the Hub, ungated)
- data
- described(publicly available datasets and an in-house multilingual document corpus, named in the release post
- code
- partial(Transformers inference, a NeMo AutoModel fine-tuning recipe made with NVIDIA and community Axolotl support
- license
- Apache-2.0(OSI
North Micro Vision's weights are released under the Apache 2.0 license and are not gated. Cohere trained it partly on its own document corpus and publishes neither that data nor its training code.
- https://docs.cohere.com/docs/models recorded 2026-09-27
Cohere's API lists north-small-translate-1-0 and north-mini-code-1-0 as the North models it serves; North Micro Vision is not an API tier
- https://huggingface.co/api/models/CohereLabs/North-Micro-Vision-Instruct?expand[]=downloads&expand[]=cardData&expand[]=gated&expand[]=createdAt&expand[]=lastModified&expand[]=siblings recorded 2026-09-27
Hub metadata for North-Micro-Vision-Instruct: gated false, model.safetensors, card license apache-2.0
- https://huggingface.co/blog/CohereLabs/meet-north-micro-vision-instruct recorded 2026-09-27
"The curriculum drew on publicly available datasets and an in-house, large-scale multilingual document corpus"
- https://huggingface.co/CohereLabs/North-Micro-Vision-Instruct/raw/main/README.md recorded 2026-09-27
"released under the Apache 2.0 license"; "Apache 2.0-licensed model weights"; fine-tuning via an "NVIDIA AutoModel recipe" and "Community-supported fine-tuning using the Axolotl framework"
Adoption
3 high confidenceHugging Face downloads over the trailing 30 days of North Micro Vision Instruct, the only North vision checkpoint.
- https://huggingface.co/api/models?author=CohereLabs&search=North&limit=1000&expand[]=downloads&expand[]=createdAt&expand[]=lastModified&expand[]=cardData&expand[]=gated recorded 2026-09-27
North-Micro-Vision-Instruct 168,910 downloads in the trailing 30 days; no other North vision checkpoint
Capability
2 high confidenceNorth Micro Vision reads documents and charts well for its size and takes several images at once. Its card claims no video understanding and rules out tool use, so it sits level with Moondream.
- https://huggingface.co/CohereLabs/North-Micro-Vision-Instruct/raw/main/README.md recorded 2026-09-27
"Inputs | Interleaved text and images"; "Broad image-understanding capabilities across VQA, captioning, grounding, OCR, charts, and documents"; "Tool calling and agentic workflows are not supported"; DocVQA_VAL 0.921, ChartQA_Test 0.808, OCRBench 0.792
Verified 2026-09-27