Nemotron Nano VL
NVIDIANemotron Nano VL is NVIDIA's vision-language model line for document intelligence, image question answering and video understanding. The current Nemotron Nano 12B v2 VL reads up to four document images at 1k by 2k resolution and samples video at up to 128 frames in a 128K context. The earlier Llama-3.1-Nemotron-Nano-VL-8B is built on Llama 3.1.
Openness
3 high confidence- weights
- open(safetensors on the Hub, ungated)
- data
- partial(Nemotron-VLM-Dataset-v2 is released under CC-BY-4.0
- code
- partial(inference snippets only
- license
- NVIDIA-Open-Model-License("Models are commercially usable")
The weights are released under the NVIDIA Open Model License, which permits commercial use. NVIDIA publishes a large vision-language training set alongside the model, but the card says the mix also used internal data, and no training code is released.
- https://huggingface.co/api/datasets/nvidia/Nemotron-VLM-Dataset-v2 recorded 2026-09-26
Nemotron-VLM-Dataset-v2: public, ungated, cc-by-4.0, visual question answering, image-text and video-text tasks
- https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-Nano-12B-v2-VL-BF16?expand[]=downloads&expand[]=cardData&expand[]=gated&expand[]=createdAt&expand[]=lastModified&expand[]=siblings recorded 2026-09-26
Hub metadata: gated false, seven safetensors shards, license_name nvidia-open-model-license
- https://huggingface.co/nvidia/NVIDIA-Nemotron-Nano-12B-v2-VL-BF16/raw/main/README.md recorded 2026-09-26
"The post-training datasets consist of a mix of internal and public datasets"; quick-start inference code only
- https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/ recorded 2026-09-26
"NVIDIA Open Model License Agreement": "Models are commercially usable. You are free to create and distribute Derivative Models."
Adoption
3 high confidenceHugging Face downloads over the trailing 30 days, summed across NVIDIA's own checkpoints of the 8B and 12B v2 models, including its FP8 and NVFP4 builds.
- https://huggingface.co/api/models?author=nvidia&search=Nano-VL&limit=100 recorded 2026-09-26
Llama-3.1-Nemotron-Nano-VL-8B-V1 76,940, its FP4-QAD build 455 and mcore build 0 downloads in the trailing 30 days
- https://huggingface.co/api/models?author=nvidia&search=v2-VL&limit=100 recorded 2026-09-26
NVIDIA-Nemotron-Nano-12B-v2-VL BF16 20,600, FP8 62,179 and NVFP4-QAD 11,691 downloads in the trailing 30 days
Capability
3 high confidenceNemotron Nano VL handles several document images at once and understands video, with strong document and chart scores. It documents no GUI operation or audio input, which places it level with MiniCPM-V rather than with the computer-use models.
- https://huggingface.co/nvidia/NVIDIA-Nemotron-Nano-12B-v2-VL-BF16/raw/main/README.md recorded 2026-09-26
"enables multi-image reasoning and video understanding"; "Input Images Supported: 4"; "Frames: 2 FPS with min of 8 frame and max of 128 frames"; MMMU 68, DocVQA 94.39, Video-MME w/o sub 65.9
Verified 2026-09-26