BAGEL
ByteDance Seed / Volcano EngineBAGEL is ByteDance Seed's unified multimodal model, a decoder-only mixture-of-transformers with 7B active and 14B total parameters that both understands images and generates and edits them. It was pretrained on interleaved text, image, video and web data and extends to free-form visual manipulation and multiview synthesis. ByteDance publishes the training code.
Openness
3 high confidence- weights
- open(safetensors on the Hub, ungated)
- data
- described(the trillions of pretraining tokens are described
- code
- open(the unified pretraining and fine-tuning pipeline (train/pretrain_unified_navit.py, scripts/train.sh))
- license
- Apache-2.0(OSI
BAGEL's weights and training code are Apache 2.0, so the pipeline can be rerun on new data. The pretraining corpus is described but not released, so the model itself cannot be rebuilt.
- https://cdn.jsdelivr.net/gh/ByteDance-Seed/Bagel@main/LICENSE recorded 2026-09-26
The repository LICENSE is the Apache License, Version 2.0
- https://cdn.jsdelivr.net/gh/ByteDance-Seed/Bagel@main/scripts/train.sh recorded 2026-09-27
The launch script for train/pretrain_unified_navit.py
- https://cdn.jsdelivr.net/gh/ByteDance-Seed/Bagel@main/TRAIN.md recorded 2026-09-27
"We provide data examples for T2I, Editing, and VLM tasks" built from FLUX.1-dev outputs, SEED-Data-Edit and LLaVA-OneVision-Data
- https://huggingface.co/api/models/ByteDance-Seed/BAGEL-7B-MoT?expand[]=downloads&expand[]=cardData&expand[]=gated&expand[]=createdAt&expand[]=lastModified&expand[]=siblings recorded 2026-09-26
Hub metadata for BAGEL-7B-MoT: gated false, ema.safetensors and ae.safetensors, card license apache-2.0
Adoption
1 high confidenceHugging Face downloads over the trailing 30 days for ByteDance's single BAGEL checkpoint. Copies fetched through the repository's own download instructions are counted only if they come from the Hub.
- https://huggingface.co/api/models?author=ByteDance-Seed&search=BAGEL&limit=100 recorded 2026-09-26
One checkpoint, BAGEL-7B-MoT, with 872 downloads in the trailing 30 days
Capability
1 medium confidenceBAGEL understands single images about as well as Qwen2.5-VL-7B and also generates and edits them. Its understanding is reported on single-image benchmarks with no document or video evaluation, so it sits level with Florence-2; the generation is recorded but does not raise it.
- https://arxiv.org/abs/2505.14683 recorded 2026-09-26
"a unified, decoder-only model pretrained on trillions of tokens curated from large-scale interleaved text, image, video, and web data"
- https://cdn.jsdelivr.net/gh/ByteDance-Seed/Bagel@main/README.md recorded 2026-09-26
Visual Understanding table: BAGEL MME 2388, MMBench 85.0, MMMU 55.3, MathVista 73.1; Qwen2.5-VL-7B 2347, 83.5, 58.6, 68.2; GenEval 0.82
Verified 2026-09-26