AI Potluck
Back to Gap Map Model components / Multimodal models

Nemotron Nano Omni

NVIDIA
open weights / Overall score: 4.7(leading)

Nemotron Nano Omni is NVIDIA's omni-modal reasoning model, a 30B-total, 3B-active mixture-of-experts that understands video, audio, images and text and answers in text. It reads video with its soundtrack, transcribes speech with word-level timestamps, parses long documents and drives computer interfaces, with tool calling in a 256K context. NVIDIA ships BF16, FP8 and NVFP4 builds.

Openness

3 high confidence
3.0
weights
open(safetensors on the Hub, ungated)
data
partial(Nemotron-Image-Training-v3 is released under CC-BY-4.0
code
closed(serving and deployment instructions only)
license
NVIDIA-Open-Model-Agreement("Works are commercially usable"

The weights are published under NVIDIA's Open Model Agreement, which permits commercial use and derivatives. NVIDIA releases one of its image training sets, but the full training corpus includes licensed and internal data and no training code is published. NVIDIA-Open-Model-Agreement restricts neither who may use it nor at what scale, so the license tier is permissive_non_osi.

Adoption

4 high confidence
4.0

Hugging Face downloads over the trailing 30 days, summed across NVIDIA's BF16, FP8 and NVFP4 builds of the model. Most of it is the quantized builds rather than the BF16 original.

Capability

5 high confidence
5.0

Nemotron Nano Omni understands audio and video natively together with images and text, and also operates computer interfaces, scoring 47.4 on OSWorld. It sits level with Qwen-Omni, with one difference: it answers in text only, where Qwen-Omni also speaks.

Verified 2026-09-27