AI Potluck
Back to Gap Map Model components / Image, video, 3D & music generation

Z-Image

Alibaba Cloud
open weights / Overall score: 3.0

Z-Image is a 6B-parameter single-stream diffusion transformer for text-to-image generation from Alibaba's Tongyi-MAI lab. Z-Image-Turbo is a distilled version that generates in eight steps and fits in 16 GB of GPU memory, with photorealistic output and bilingual English and Chinese text rendering; the undistilled Z-Image base model supports full classifier-free guidance for fine-tuning.

Openness

3 high confidence
3.0
weights
open(Z-Image-Turbo and Z-Image download from Hugging Face without a gate)
data
closed(the cards and README neither release nor describe the training data)
code
partial(the repository holds inference code only
license
Apache-2.0(OSI

Both Z-Image checkpoints are released under Apache 2.0, with no gate on either. The first-party repository holds inference code only, and the training data is not published.

Adoption

3 high confidence
3.0

Adoption is measured as Hugging Face downloads of the two checkpoints, nearly all of them for Z-Image-Turbo.

Capability

3 medium confidence
3.0

Z-Image Turbo places in the upper half of the text-to-image arena, a strong result for a model small enough for a consumer GPU. It trails FLUX.2 [dev] there and sits in the same group of open models.

Verified 2026-09-26