AI Potluck
Back to Gap Map Model components / Image, video, 3D & music generation

Mochi

Genmo
open weights / Overall score: 1.6

Mochi 1 is Genmo's 10B-parameter text-to-video diffusion model, built on its Asymmetric Diffusion Transformer architecture and trained from scratch. It generates 480p clips and is tuned for photorealistic rather than animated styles. Genmo publishes an inference harness with context parallelism and a LoRA fine-tuning trainer that runs on a single 80 GB GPU.

Openness

3 high confidence
3.0
weights
open(mochi-1-preview downloads from Hugging Face without a gate)
data
closed(the README says only that the model was trained from scratch
code
partial(inference harness and a LoRA fine-tuning trainer
license
Apache-2.0(OSI

Mochi 1 is released under Apache 2.0, with no gate on the checkpoint. Genmo publishes inference and LoRA fine-tuning code but not the pretraining pipeline or data.

Adoption

1 high confidence
1.0

Adoption is measured as Hugging Face downloads of the mochi-1-preview checkpoint.

Capability

2 medium confidence
2.0

Mochi 1 places in the lower half of the video arena, below the downloadable Wan 2.2 and HunyuanVideo. It sits a step below LTX-2.5, which places in the upper half.

Verified 2026-09-26