LatentSync
ByteDance Seed / Volcano EngineLatentSync is ByteDance's audio-conditioned latent diffusion model for lip-sync: given a video of a face and a speech track, it regenerates the mouth region so the lips match the audio, supervised by a SyncNet model released with it. The current release, LatentSync 1.6, is trained on 512x512 video.
Lip-sync had no head product: LivePortrait animates a portrait from a driving video, while LatentSync edits an existing video to match new speech. The code is Apache-2.0; the Hub checkpoints carry only an openrail++ tag in their card metadata, with no license file.
Openness
2 medium confidence- weights
- open(LatentSync 1.6 and 1.5 download from Hugging Face without a gate)
- data
- described(the training videos are described (512x512 in 1.6
- code
- open(U-Net and SyncNet training scripts and the data processing pipeline are published under Apache-2.0)
- license
- CreativeML-OpenRAIL++-M(the checkpoints, declared as openrail++ in the Hub card metadata
LatentSync's code, training scripts included, is Apache 2.0, and its own weights carry an OpenRAIL++ license that restricts some uses but not commercial use. Lip-syncing a video runs every frame through a face detector built on InsightFace models licensed for non-commercial research only, and the training videos are not released.
- https://huggingface.co/api/models/ByteDance/LatentSync-1.6?expand[]=downloads&expand[]=likes&expand[]=cardData&expand[]=gated&expand[]=lastModified&expand[]=createdAt recorded 2026-09-27
Hugging Face API record for ByteDance/LatentSync-1.6: cardData license openrail++, gated false, created 2025-06-11.
- https://huggingface.co/api/models/ByteDance/LatentSync-1.6/tree/main recorded 2026-09-27
File tree: latentsync_unet.pt and stable_syncnet.pt at the root, with auxiliary/ and whisper/ folders; no license file.
- https://huggingface.co/api/models/ByteDance/LatentSync-1.6/tree/main/auxiliary recorded 2026-09-27
auxiliary/ tree: i3d_torchscript.pt, koniq_pretrained.pkl, sfd_face.pth, syncnet_v2.model, vgg16-397923af.pth, vit_g_hybrid_pt_1200e_ssv2_ft.pth.
- https://huggingface.co/ByteDance/LatentSync-1.6/raw/main/README.md recorded 2026-09-27
Card: "the only modification was upgrading the training dataset to 512 x 512 videos"; no dataset is released.
- https://raw.githubusercontent.com/bytedance/LatentSync/main/latentsync/utils/image_processor.py recorded 2026-09-27
ImageProcessor imports FaceDetector from .face_detector and builds self.face_detector = FaceDetector(device=device); the lip-sync pipeline's affine_transform_video calls image_processor.affine_transform on every frame.
- https://ungh.cc/repos/bytedance/LatentSync/files/main recorded 2026-09-27
Repository file list: scripts/train_unet.py, scripts/train_syncnet.py and preprocess/data_processing_pipeline.py.
- https://ungh.cc/repos/bytedance/LatentSync/files/main/latentsync/utils/face_detector.py recorded 2026-09-27
face_detector.py: "from insightface.app import FaceAnalysis".
- https://ungh.cc/repos/bytedance/LatentSync/files/main/LICENSE recorded 2026-09-27
Code LICENSE body: "Apache License / Version 2.0, January 2004".
- https://ungh.cc/repos/bytedance/LatentSync/files/main/README.md recorded 2026-09-27
README: "2025/06/11: We released LatentSync 1.6", and an open-source plan that checks off inference code, checkpoints, "Data processing pipeline" and training code.
- https://ungh.cc/repos/bytedance/LatentSync/files/main/requirements.txt recorded 2026-09-27
requirements.txt: "insightface==0.7.3".
- https://ungh.cc/repos/deepinsight/insightface/files/master/README.md recorded 2026-09-27
InsightFace README: auto-downloading models with the python library follow "the above license policy(which is for non-commercial research purposes only)".
Adoption
3 high confidenceAdoption is measured as Hugging Face downloads of the LatentSync checkpoints, led by version 1.6.
- https://huggingface.co/api/models?author=ByteDance&search=LatentSync&sort=downloads&direction=-1&limit=100 recorded 2026-09-27
Listing sorted by downloads: LatentSync-1.6 199929, LatentSync-1.5 75382, LatentSync 0.
Capability
1 medium confidenceLatentSync re-renders a speaker's mouth to match new audio rather than generating video from a prompt, and no public arena ranks lip-sync. It sits level with LivePortrait, the other face-animation model here.
- https://artificialanalysis.ai/video/leaderboard/text-to-video recorded 2026-09-27
The text-to-video leaderboard lists no LatentSync entry.
- https://arxiv.org/abs/2412.09262 recorded 2026-09-27
Paper title: "LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision".
Verified 2026-09-27