MMAudio
hkchengrexMMAudio generates sound effects and ambient audio synchronized to a video clip, a text prompt, or both, trained jointly on audio-visual and audio-text data. The model is small (157M parameters in the paper) and produces an eight-second clip in about a second. It came out of the University of Illinois Urbana-Champaign with Sony AI and was presented at CVPR 2025.
Gap fill: video-to-audio (Foley and sound-effect) generation had no head product; the music models here generate songs and instrumentals from text, not sound synchronized to picture. MMAudio is the most-starred of the four seeded sound-effect rows (ThinkSound, HunyuanVideo-Foley and FoleyCrafter stay in the registry). Its adoption signal is thin: the Hub reports zero downloads for a repository that holds .pth files without a config.json, which the Hub's download counter does not count. The code is MIT; the checkpoints are CC-BY-NC 4.0.
Openness
2 high confidence- weights
- open(the small, medium and large checkpoints download from hkchengrex/MMAudio on Hugging Face without a gate)
- data
- described(trained on public sets named in the README and TRAINING.md (AudioSet, Freesound, VGGSound, AudioCaps, WavCaps, Clotho), which the authors say they cannot redistribute)
- code
- open(train.py, the training configs and latent-extraction scripts are published under MIT)
- license
- CC-BY-NC-4.0(the checkpoints, "for non-commercial purposes only"
MMAudio's weights may be used for non-commercial purposes only. The training code is published under MIT, and the public datasets it was trained on are named, though not redistributed.
- https://huggingface.co/api/models/hkchengrex/MMAudio?expand[]=downloads&expand[]=likes&expand[]=cardData&expand[]=gated&expand[]=lastModified&expand[]=createdAt&expand[]=siblings recorded 2026-09-27
Hugging Face API record: cardData license cc-by-nc-4.0, gated false; files weights/mmaudio_large_44k_v2.pth and the other checkpoints.
- https://huggingface.co/hkchengrex/MMAudio/raw/main/README.md recorded 2026-09-27
Card: "Please note that the checkpoints are licensed under CC BY-NC 4.0 and can be used for non-commercial purposes only."
- https://ungh.cc/repos/hkchengrex/MMAudio/files/main recorded 2026-09-27
Repository file list: train.py, config/train_config.yaml and training/ scripts.
- https://ungh.cc/repos/hkchengrex/MMAudio/files/main/docs/TRAINING.md recorded 2026-09-27
TRAINING.md: "We used VGGSound, AudioCaps, WavCaps, and Clotho" and "We cannot redistribute the datasets for copyright reasons".
- https://ungh.cc/repos/hkchengrex/MMAudio/files/main/LICENSE recorded 2026-09-27
Code LICENSE body: "MIT License / Copyright (c) 2024 Sony Research Inc."
- https://ungh.cc/repos/hkchengrex/MMAudio/files/main/README.md recorded 2026-09-27
README: "The checkpoints are released on Hugging Face under the CC-BY-NC 4.0 license"; "MMAudio was trained on several datasets, including AudioSet, Freesound, VGGSound, AudioCaps, and WavCaps".
Adoption
1 low confidenceThe Hub reports zero downloads for the checkpoint repository, which holds only .pth files and no config.json, so its counter likely misses real use. The level is the measured floor, not a claim that nobody uses the model.
- https://huggingface.co/api/models/hkchengrex/MMAudio?expand[]=downloads&expand[]=likes&expand[]=cardData&expand[]=gated&expand[]=lastModified&expand[]=createdAt&expand[]=siblings recorded 2026-09-27
Hugging Face API record for hkchengrex/MMAudio: downloads 0, likes 131; siblings list no config.json.
- https://ungh.cc/repos/hkchengrex/MMAudio recorded 2026-09-27
Repository record: stars 2267, pushedAt 2026-02-23.
Capability
1 medium confidenceMMAudio generates sound effects to match a video rather than music, and no public arena ranks that task. It sits below MusicGen, which at least places on the music arena.
- https://artificialanalysis.ai/music/leaderboard/instrumental recorded 2026-09-27
The 19-entry instrumental board lists no MMAudio entry; "MusicGen" is on it.
- https://arxiv.org/abs/2412.15322 recorded 2026-09-27
Abstract: "MMAudio achieves new video-to-audio state-of-the-art among public models ... and just 157M parameters."
Verified 2026-09-27