AI Potluck
Back to Gap Map Model components / Multimodal models

BAGEL

ByteDance Seed / Volcano Engine
open weights / Overall score: 1.0

BAGEL is ByteDance Seed's unified multimodal model, a decoder-only mixture-of-transformers with 7B active and 14B total parameters that both understands images and generates and edits them. It was pretrained on interleaved text, image, video and web data and extends to free-form visual manipulation and multiview synthesis. ByteDance publishes the training code.

Openness

3 high confidence
3.0
weights
open(safetensors on the Hub, ungated)
data
described(the trillions of pretraining tokens are described
code
open(the unified pretraining and fine-tuning pipeline (train/pretrain_unified_navit.py, scripts/train.sh))
license
Apache-2.0(OSI

BAGEL's weights and training code are Apache 2.0, so the pipeline can be rerun on new data. The pretraining corpus is described but not released, so the model itself cannot be rebuilt.

Adoption

1 high confidence
1.0

Hugging Face downloads over the trailing 30 days for ByteDance's single BAGEL checkpoint. Copies fetched through the repository's own download instructions are counted only if they come from the Hub.

Capability

1 medium confidence
1.0

BAGEL understands single images about as well as Qwen2.5-VL-7B and also generates and edits them. Its understanding is reported on single-image benchmarks with no document or video evaluation, so it sits level with Florence-2; the generation is recorded but does not raise it.

Verified 2026-09-26