AI Potluck
Back to Gap Map Model components / Robotics & embodied AI

RDT (Robotics Diffusion Transformer)

Tsinghua Machine Learning Group
open weights / Overall score: 2.8

RDT is Tsinghua University's line of bimanual robot foundation models. RDT-1B is a one-billion-parameter diffusion transformer pretrained on more than a million multi-robot episodes; RDT2 adapts Qwen2.5-VL-7B into a vision-language-action model trained on over ten thousand hours of data from a handheld gripper, and deploys zero-shot on bimanual UR5e and Franka arms it was not trained on.

Openness

3 high confidence
3.0
weights
open(RDT2 and RDT-1B checkpoints on the Hub)
data
described(RDT2's 10,000+ hours of UMI data described, not released)
code
open(training and fine-tuning scripts)
license
Apache-2.0(RDT2)+MIT(RDT-1B)

Weights and training code are released under Apache-2.0 and MIT. The first release also published its fine-tuning data, but the ten thousand hours of gripper recordings that RDT2 learned from are described and not published.

Adoption

1 medium confidence
1.0

Hugging Face downloads summed over the RDT2 and RDT-1B checkpoints.

Capability

4 low confidence
4.0

RDT learns from many robots and is built around two-armed manipulation, and RDT2 claims to run on bimanual arms it never saw in training, which few open policies attempt. Its tasks are simple ones such as picking, placing and wiping.

Verified 2026-09-26