AI Potluck
Back to Gap Map Infrastructure / Classic ML & computer vision

Depth Anything

ByteDance Seed / Volcano Engine
restricted / Overall score: 4.0(strong)

Depth-estimation model line from ByteDance Seed. Depth Anything 3, the current release, predicts consistent depth and camera pose from one image or many, with or without known poses, using a plain DINO-style transformer, and ships any-view, metric-depth, monocular and nested checkpoints from Small to Giant. Depth Anything V1 and V2 remain widely used for single-image depth.

Openness

2 high confidence
2.0
weights
open(every DA3 checkpoint on the Hub, ungated)
data
documented-not-released(stated to be public academic datasets only
code
partial(inference, CLI and benchmark code
license
Apache-2.0(code, and the Small, Base, Metric and Mono checkpoints)+CC-BY-NC-4.0(the Large, Giant and Nested any-view checkpoints)

The code and the small, base, metric and monocular models are Apache-2.0, but the large, giant and nested any-view models, the line's flagships, are CC BY-NC 4.0 and may not be used commercially. The map treats the most restrictive license among a release's distributed checkpoints as the release's. One Hub card, DA3-LARGE-1.1, is tagged Apache-2.0 while the repository's license table lists it as CC BY-NC; the giant models settle the reading either way.

Adoption

4 high confidence
4.0

Hugging Face downloads over the trailing 30 days for the six declared Depth Anything 3 checkpoints. The older V2 and V1 checkpoints, Depth-Anything-V2-Small above all, are downloaded more than the current release; they would raise the total without changing its order of magnitude.

Capability

4 medium confidence
4.0

Depth Anything gives a depth map, and with several views a camera pose, for any image the user supplies without training, and depth has no label set that a new task would require retraining for. It covers the geometry tasks only, where Ultralytics spans detection, segmentation and depth together.

  • https://raw.githubusercontent.com/ByteDance-Seed/Depth-Anything-3/main/README.md recorded 2026-09-27

    README: DA3 "predicts spatially consistent geometry from arbitrary visual inputs, with or without known camera poses"; capabilities "Monocular Depth Estimation", "Pose-Conditioned Depth Estimation", "Camera Pose Estimation"; "DA3 Metric Series ... fine-tuned for metric depth estimation in monocular settings".

Verified 2026-09-27