AI Potluck
Back to Gap Map Infrastructure / Classic ML & computer vision

RT-DETR

lyuwenyu
open weights / Overall score: 3.6

Real-time end-to-end object detector family built on the DETR transformer design, which removes the non-maximum-suppression step that YOLO-style detectors need. The line began with RT-DETR and RT-DETRv2, published from the lyuwenyu repository; RT-DETRv4, the newest member, distills features from a DINOv3 vision foundation model into small detectors. Checkpoints are trained on COCO, some pretrained on Objects365, and full training code ships for each version.

Openness

4 medium confidence
4.0
weights
open(RT-DETRv4 S to X checkpoints linked from its README
data
documented-not-released(COCO 2017 is public, but v4 distills from a DINOv3 ViT-B/16 teacher trained on the unreleased LVD-1689M corpus)
code
open(training code and configs for v1, v2 and v4)
license
Apache-2.0(repository LICENSE for v1, v2 and v4

Every version is Apache-2.0, code and weights, with the full training code and configs published. The training data cannot be fully reproduced, though: RT-DETRv4's recipe distills features from a DINOv3 ViT-B/16 teacher, and Meta has not released that teacher's LVD-1689M training images, so COCO being public does not make the detector's full training data available. The teacher itself must be obtained separately under Meta's DINOv3 License.

Adoption

4 high confidence
4.0

Hugging Face downloads over the trailing 30 days for the two most-used checkpoints, which are the ones declared; the other PekingU checkpoints add comparatively little. RT-DETRv4 checkpoints are downloaded from a file-sharing link that publishes no counts.

Capability

3 high confidence
3.0

RT-DETR and RF-DETR are the same kind of product, real-time transformer detectors pretrained on COCO's 80 classes, so detecting anything else still takes a training run on the user's own labels. RT-DETRv4 reaches 57.0 COCO AP at its largest size.

Verified 2026-09-27