MiMo-V2.5-Pro
XiaomiXiaomi's open-weights MoE model (released April 2026): ~1.02T total / 42B active params (70 layers, 8 experts/token), hybrid attention interleaving sliding-window and global attention (6:1), 3-layer Multi-Token Prediction, and a 1M-token context. Built for demanding agentic, complex software-engineering, and long-horizon tasks.
Open weights, MIT per the HF model card (an older MiMo GitHub repo lists Apache-2.0; the model card is authoritative for this release). ~1.02T/42B MoE, 1M context. ~66k HF downloads/month plus third-party quantizations.
Openness
3 medium confidence- weights
- open(MIT per HF card)
- data
- closed
- code
- open
- license
- MIT(HF card
Open weights under a permissive license (MIT per the model card); no open training data -> open_weights.
- https://huggingface.co/XiaomiMiMo/MiMo-V2.5-Pro recorded 2026-06-18
MIT, 1.02T/42B MoE, 1M context, benchmarks
Adoption
4 medium confidence~66,027 HF downloads last month, 651 likes, plus third-party quantizations (unsloth GGUF, MLX, vLLM).
- https://huggingface.co/XiaomiMiMo/MiMo-V2.5-Pro recorded 2026-06-18
66,027 downloads last month, 651 likes
Capability
4 medium confidenceFrontier open-weights model, efficiency-focused; official benchmarks, Arena rank mirror-reported.
- https://mimo.xiaomi.com/mimo-v2-5-pro recorded 2026-06-18
official architecture + benchmarks, Apr 2026 release
Unchanged since 2026-07-30 (last edited, not re-checked)