AI Potluck
Back to Gap Map Model components / Image, video, 3D & music generation

Stable Video Diffusion

Stability AI
open weights / Overall score: 1.8

Stable Video Diffusion is Stability AI's latent diffusion model that turns a single still image into a short video clip. The original SVD generates 14 frames and SVD-XT 25 frames at 576x1024; SVD-XT 1.1, released in February 2024, is a fine-tune of SVD-XT for more consistent output.

Stable Video 3D (SV3D) and Stable Video 4D (SV4D, SV4D 2.0 of May 2025) are built on SVD but generate orbital and multi-view video of an object rather than a clip from an image, so they are not part of this entry; the category already treats Stability AI's 3D models as separate products. The generative-models repository is shared with Stable Diffusion and publishes SVD inference only. Dormant: no SVD release since February 2024, but the two 2023 checkpoints still draw about 330K downloads a month.

Openness

3 medium confidence
3.0
weights
open(SVD and SVD-XT download without a gate
data
described(the SVD paper and cards describe the curated training video set and its filtering
code
partial(Stability-AI/generative-models publishes SVD inference configs and sampling scripts under MIT
license
Stability-AI-Community-License(the LICENSE.md in every SVD repository: free below USD 1M annual revenue, after which an enterprise license is required

Stable Video Diffusion is free to use, commercially included, for organizations under USD 1M in annual revenue, with a paid license above that; the newest checkpoint's download form still asks for non-commercial use. Stability publishes inference code only, and the training videos are described but not released. Stability-AI-Community-License allows commercial use only within a bound, so the license tier is use_bounded.

Adoption

3 high confidence
3.0

Adoption is measured as Hugging Face downloads across the SVD checkpoints, almost all of them the original SVD-XT and SVD rather than the newer 1.1.

Capability

1 medium confidence
1.0

Stable Video Diffusion generates only from an image and has no entry on either public video arena. It sits below Mochi 1, which places on the text-to-video board.

Verified 2026-09-27