AI Potluck
Back to Gap Map Model components / Multimodal models

MiniCPM-o

OpenBMB
open weights / Overall score: 4.7(leading)

MiniCPM-o is OpenBMB's omni-modal model line for live, full-duplex interaction on phones and other devices. MiniCPM-o 4.5, a 9B model built from SigLIP2, Whisper-medium, CosyVoice2 and Qwen3-8B, watches video and listens to audio at the same time as it answers in text and speech, and can clone a voice from a short reference clip. It shares a repository with MiniCPM-V.

Openness

3 high confidence
3.0
weights
open(safetensors and the speech decoder on the Hub, ungated)
data
described(the 4.5 card mentions curated post-training data and professional voice-actor recordings
code
partial(the shared finetune/ scripts name MiniCPM-o 2.6
license
Apache-2.0(OSI

MiniCPM-o 4.5 is Apache 2.0 and ships its speech decoder with the weights. OpenBMB describes its post-training data and voice recordings but publishes neither, and offers fine-tuning guides rather than its own training recipe.

Adoption

4 high confidence
4.0

Hugging Face downloads over the trailing 30 days, summed across OpenBMB's MiniCPM-o 2.6 and 4.5 checkpoints, including its GGUF, AWQ and int4 builds. The search also returns OpenBMB's text-only MiniCPM models, which are excluded.

Capability

5 high confidence
5.0

MiniCPM-o sees, listens and speaks at the same time, taking continuous video and audio and answering in text and speech, and its Video-MME score matches Qwen3-Omni's. It sits level with Qwen-Omni, with speech in two languages rather than ten.

Verified 2026-09-27