AI Potluck
Back to Gap Map Model components / Image, video, 3D & music generation

Open-Sora

HPC-AI Tech
open weights / Overall score: 1.0

Open-Sora is HPC-AI Tech's open video-generation model family. The current release, Open-Sora 2.0, is an 11B-parameter model for text-to-video and image-to-video at up to 768px, trained for about USD 200K with the checkpoints and training code both published. Its own configs generate text-to-video directly, and an optional text-to-image-to-video pipeline uses FLUX.1 [dev] for the first frame.

Gap fill: the video models here are either closed or ship weights with at most inference or fine-tuning code; Open-Sora is the one video line that publishes its Apache-2.0 checkpoints together with the pretraining code. The training data itself is not released. Adoption is thin (under 4K Hub downloads a month across the family) despite 29.8K GitHub stars. The Open-Sora-v2 repository also redistributes, under their own licenses, the HunyuanVideo VAE every generation config loads and the FLUX.1 [dev] weights used only by the optional text-to-image-to-video pipeline. HPC-AI Tech's hosted Video Ocean product says it runs on "a superior model"; no newer Open-Sora release is named, so 2.0 is taken as current.

Openness

3 medium confidence
3.0
weights
open(Open_Sora_v2.safetensors downloads from hpcai-tech/Open-Sora-v2 without a gate
data
described(the training data is described in the technical report
code
open(the pretraining scripts (scripts/diffusion/train.py) and docs/train.md are published
license
Apache-2.0(the Open-Sora 2.0 weights and the repository)+Tencent-Hunyuan-Community-License(the HunyuanVideo VAE that every text-to-video and image-to-video config loads, redistributed in the Open-Sora-v2 repository, and code the LICENSE bundles
optional-pipeline
FLUX.1-dev-Non-Commercial-License(FLUX.1 [dev] weights sit in the same Hugging Face repository and the README's download command fetches them, but only the optional text-to-image-to-video pipeline uses them

Open-Sora 2.0's own weights and training code are Apache 2.0, but every text-to-video and image-to-video run loads a HunyuanVideo VAE whose license excludes the EU, UK and South Korea, which caps it at use_bounded. FLUX.1 [dev] weights under a non-commercial license sit in the same repository for an optional text-to-image-to-video pipeline; the direct text-to-video path does not load them, so they do not govern. The training data is described, not released.

Adoption

1 high confidence
1.0

Adoption is measured as Hugging Face downloads across the Open-Sora checkpoints, led by the 1.2 VAE and the 2.0 model. The project's GitHub following is far larger than its measured use.

Capability

1 medium confidence
1.0

Open-Sora does not appear on the public video arena, where Mochi 1 places in the lower half. Its own report compares it with HunyuanVideo, but the map reads one public board per modality.

Verified 2026-09-27