AI Potluck
Back to Gap Map Model components / Multimodal models

Nemotron Nano VL

NVIDIA
open weights / Overall score: 3.0

Nemotron Nano VL is NVIDIA's vision-language model line for document intelligence, image question answering and video understanding. The current Nemotron Nano 12B v2 VL reads up to four document images at 1k by 2k resolution and samples video at up to 128 frames in a 128K context. The earlier Llama-3.1-Nemotron-Nano-VL-8B is built on Llama 3.1.

Openness

3 high confidence
3.0
weights
open(safetensors on the Hub, ungated)
data
partial(Nemotron-VLM-Dataset-v2 is released under CC-BY-4.0
code
partial(inference snippets only
license
NVIDIA-Open-Model-License("Models are commercially usable")

The weights are released under the NVIDIA Open Model License, which permits commercial use. NVIDIA publishes a large vision-language training set alongside the model, but the card says the mix also used internal data, and no training code is released.

Adoption

3 high confidence
3.0

Hugging Face downloads over the trailing 30 days, summed across NVIDIA's own checkpoints of the 8B and 12B v2 models, including its FP8 and NVFP4 builds.

Capability

3 high confidence
3.0

Nemotron Nano VL handles several document images at once and understands video, with strong document and chart scores. It documents no GUI operation or audio input, which places it level with MiniCPM-V rather than with the computer-use models.

Verified 2026-09-26