AI Potluck
Back to Gap Map Model components / Multimodal models

Janus

DeepSeek
open weights / Overall score: 1.3

Janus is DeepSeek's family of unified models that both understand images and generate them from text in one autoregressive transformer. It separates visual encoding for the two jobs, SigLIP for understanding and a discrete image tokenizer for generation. Janus-Pro comes in 1B and 7B sizes on DeepSeek-LLM bases, and JanusFlow swaps in rectified flow for generation.

Openness

3 high confidence
3.0
weights
open(PyTorch .bin weights on the Hub, ungated)
data
described(the Janus-Pro paper names the SFT additions and about 72 million synthetic aesthetic samples
code
partial(inference, generation and demo code
license
DeepSeek-Model-License(the DeepSeek License Agreement 1.0, word for word the DeepSeek-LLM LICENSE-MODEL

Janus-Pro is released under DeepSeek's model license, which allows commercial use but forbids a list of harmful uses and requires passing those limits on. Only inference code is published, and the training data is described but not released.

Adoption

2 high confidence
2.0

Hugging Face downloads over the trailing 30 days, summed across DeepSeek's four Janus checkpoints.

Capability

1 high confidence
1.0

Janus-Pro answers questions about a single low-resolution image and also generates images from text. Its understanding side takes one image with no document or video focus, so it sits level with Florence-2; the image generation is recorded but does not raise it.

Verified 2026-09-26