AI Potluck
Back to Gap Map Model components / Multimodal models

UI-TARS

ByteDance Seed / Volcano Engine
open weights / Overall score: 3.7

UI-TARS is ByteDance Seed's GUI agent model, which takes screenshots as its only input and returns mouse, keyboard and touch actions for desktop, browser and phone interfaces. The released checkpoints are the 2B, 7B and 72B SFT and DPO models and UI-TARS-1.5-7B, built on a Qwen2.5-VL base. The UI-TARS-desktop application that drives it is a separate project.

The strongest UI-TARS-1.5 model and UI-TARS-2 are not downloadable; UI-TARS-1.5 is offered to researchers on request, and the scores here are read from the released 7B.

Openness

3 high confidence
3.0
weights
open(safetensors on the Hub, ungated)
data
described(the paper and README describe large-scale GUI screenshots and action traces
code
partial(action parser, prompt templates and inference tests
license
Apache-2.0(OSI

The released checkpoints are Apache 2.0 and the repository carries the prompt templates and action parser needed to run them. The GUI trajectory data and the training pipeline are not published.

Adoption

3 high confidence
3.0

Hugging Face downloads over the trailing 30 days, summed across ByteDance Seed's six UI-TARS checkpoints; nearly all of it is UI-TARS-1.5-7B. Use through the UI-TARS-desktop application is counted only where it downloads these weights.

Capability

4 high confidence
4.0

UI-TARS acts on what it sees, turning screenshots into clicks, keystrokes and phone gestures across desktop, web and Android. It sits a step below Qwen-Omni because it understands only images, with no audio or video input.

Verified 2026-09-26