AI Potluck
Back to Gap Map Model components / Image, video, 3D & music generation

MMAudio

hkchengrex
restricted / Overall score: 1.0

MMAudio generates sound effects and ambient audio synchronized to a video clip, a text prompt, or both, trained jointly on audio-visual and audio-text data. The model is small (157M parameters in the paper) and produces an eight-second clip in about a second. It came out of the University of Illinois Urbana-Champaign with Sony AI and was presented at CVPR 2025.

Gap fill: video-to-audio (Foley and sound-effect) generation had no head product; the music models here generate songs and instrumentals from text, not sound synchronized to picture. MMAudio is the most-starred of the four seeded sound-effect rows (ThinkSound, HunyuanVideo-Foley and FoleyCrafter stay in the registry). Its adoption signal is thin: the Hub reports zero downloads for a repository that holds .pth files without a config.json, which the Hub's download counter does not count. The code is MIT; the checkpoints are CC-BY-NC 4.0.

Openness

2 high confidence
2.0
weights
open(the small, medium and large checkpoints download from hkchengrex/MMAudio on Hugging Face without a gate)
data
described(trained on public sets named in the README and TRAINING.md (AudioSet, Freesound, VGGSound, AudioCaps, WavCaps, Clotho), which the authors say they cannot redistribute)
code
open(train.py, the training configs and latent-extraction scripts are published under MIT)
license
CC-BY-NC-4.0(the checkpoints, "for non-commercial purposes only"

MMAudio's weights may be used for non-commercial purposes only. The training code is published under MIT, and the public datasets it was trained on are named, though not redistributed.

Adoption

1 low confidence
1.0

The Hub reports zero downloads for the checkpoint repository, which holds only .pth files and no config.json, so its counter likely misses real use. The level is the measured floor, not a claim that nobody uses the model.

Capability

1 medium confidence
1.0

MMAudio generates sound effects to match a video rather than music, and no public arena ranks that task. It sits below MusicGen, which at least places on the music arena.

Verified 2026-09-27