AI Potluck
Back to Gap Map Infrastructure / Classic ML & computer vision

Segment Anything (SAM)

Meta
open weights / Overall score: 4.0(strong)

Meta's promptable segmentation model line for images and video. SAM 3, the current generation, detects, segments and tracks every instance of a concept named by a short text phrase or shown by an example, as well as objects picked with points, boxes and masks; SAM 3.1 adds faster multi-object tracking. The earlier SAM and SAM 2 releases, Apache-licensed, segment from visual prompts only.

Openness

3 medium confidence
3.0
weights
open(SAM 3 and SAM 3.1 checkpoints on the Hub behind a manual access request)
data
documented-not-released(the SA-Co training data from Meta's data engine is described
code
partial(inference and fine-tuning code
license
SAM-License(code and weights

SAM 3 and SAM 3.1 come under Meta's SAM License, which permits commercial use but bars military, weapons and trade-controlled uses and lets Meta change the terms, and the checkpoints need an approved access request. Meta publishes fine-tuning code and evaluation benchmarks but not the training data or pretraining code. SAM and SAM 2 were Apache-2.0, but the current release governs. The SAM License caps neither who may use the model nor at what scale, so its tier is permissive_non_osi.

Adoption

4 high confidence
4.0

Hugging Face downloads over the trailing 30 days for the two SAM 3 releases, which are the ones declared. The superseded SAM 2 and original SAM checkpoints add more and are not declared. The sam3 PyPI package, which installs the code, is declared but not counted.

Capability

4 high confidence
4.0

SAM 3 segments whatever the user names or points at in their own images and video with no training, the same no-training capability that places Ultralytics here. Ultralytics sits here partly because it bundles SAM 3; the model line itself does one family of tasks, segmentation and tracking, where Ultralytics spans many.

  • https://raw.githubusercontent.com/facebookresearch/sam3/main/README.md recorded 2026-09-27

    README: "SAM 3 is a unified foundation model for promptable segmentation in images and videos. It can detect, segment, and track objects using text or visual prompts such as points, boxes, and masks." and "SAM 3 introduces the ability to exhaustively segment all instances of an open-vocabulary concept specified by a short text phrase or exemplars".

Verified 2026-09-27