AI Potluck
Back to Gap Map Model components / Image, video, 3D & music generation

LatentSync

ByteDance Seed / Volcano Engine
restricted / Overall score: 1.8

LatentSync is ByteDance's audio-conditioned latent diffusion model for lip-sync: given a video of a face and a speech track, it regenerates the mouth region so the lips match the audio, supervised by a SyncNet model released with it. The current release, LatentSync 1.6, is trained on 512x512 video.

Lip-sync had no head product: LivePortrait animates a portrait from a driving video, while LatentSync edits an existing video to match new speech. The code is Apache-2.0; the Hub checkpoints carry only an openrail++ tag in their card metadata, with no license file.

Openness

2 medium confidence
2.0
weights
open(LatentSync 1.6 and 1.5 download from Hugging Face without a gate)
data
described(the training videos are described (512x512 in 1.6
code
open(U-Net and SyncNet training scripts and the data processing pipeline are published under Apache-2.0)
license
CreativeML-OpenRAIL++-M(the checkpoints, declared as openrail++ in the Hub card metadata

LatentSync's code, training scripts included, is Apache 2.0, and its own weights carry an OpenRAIL++ license that restricts some uses but not commercial use. Lip-syncing a video runs every frame through a face detector built on InsightFace models licensed for non-commercial research only, and the training videos are not released.

Adoption

3 high confidence
3.0

Adoption is measured as Hugging Face downloads of the LatentSync checkpoints, led by version 1.6.

Capability

1 medium confidence
1.0

LatentSync re-renders a speaker's mouth to match new audio rather than generating video from a prompt, and no public arena ranks lip-sync. It sits level with LivePortrait, the other face-animation model here.

Verified 2026-09-27