AI Potluck
Back to Gap Map Infrastructure / Classic ML & computer vision

DINO (DINOv2, DINOv3)

Meta
open weights / Overall score: 4.0(strong)

Meta's self-supervised vision backbones, trained without labels to produce general image features. DINOv3, the current release, spans ViT models from 21 million to 7 billion parameters and ConvNeXt models, trained on 1.7 billion curated web images, with a satellite-imagery variant. Meta also publishes heads that turn the features into classification, depth, detection and segmentation, and dino.txt for text-prompted zero-shot tasks. DINOv2 was the Apache-licensed predecessor.

Openness

4 medium confidence
4.0
weights
open(every DINOv3 checkpoint on the Hub behind a manual access request)
data
documented-not-released(LVD-1689M, curated from public Instagram posts, and SAT-493M satellite imagery
code
open(full pretraining pipeline, configs and distillation recipes)
license
DINOv3-License(code and weights

DINOv3's code and weights come under Meta's own DINOv3 License, which permits commercial use but bars military, weapons and trade-controlled uses and lets Meta change the terms; downloads need an approved access request. The full pretraining code is published, but the 1.7-billion-image training set drawn from Instagram is not. The earlier DINOv2 was Apache-2.0, but the current release governs. The DINOv3 License caps neither who may use the model nor at what scale, so its tier is permissive_non_osi.

Adoption

4 high confidence
4.0

Hugging Face downloads over the trailing 30 days for the three most-used DINOv3 checkpoints, which are the ones declared. The other DINOv3 checkpoints and the older, more downloaded DINOv2 checkpoints would raise the total without changing its order of magnitude.

Capability

4 medium confidence
4.0

DINOv3 is best known as a backbone that other models are built on, but Meta also ships ready heads, and its dino.txt variant classifies and segments whatever labels the user types with no training, the same kind of open-vocabulary capability that places Ultralytics here. Its detector and segmentor heads cover fixed benchmark label sets.

  • https://raw.githubusercontent.com/facebookresearch/dinov3/main/README.md recorded 2026-09-27

    README sections: "Pretrained heads - Image classification", "Pretrained heads - Depther trained on SYNTHMIX dataset", "Pretrained heads - Detector trained on COCO2017 dataset", "Pretrained heads - Segmentor trained on ADE20K dataset", "Pretrained heads - Zero-shot tasks with dino.txt"; notebook "Zero-shot segmentation with DINOv3-based dino.txt: compute the open-vocabulary segmentation results".

Verified 2026-09-27