MiniCPM-o
OpenBMBMiniCPM-o is OpenBMB's omni-modal model line for live, full-duplex interaction on phones and other devices. MiniCPM-o 4.5, a 9B model built from SigLIP2, Whisper-medium, CosyVoice2 and Qwen3-8B, watches video and listens to audio at the same time as it answers in text and speech, and can clone a voice from a short reference clip. It shares a repository with MiniCPM-V.
Openness
3 high confidence- weights
- open(safetensors and the speech decoder on the Hub, ungated)
- data
- described(the 4.5 card mentions curated post-training data and professional voice-actor recordings
- code
- partial(the shared finetune/ scripts name MiniCPM-o 2.6
- license
- Apache-2.0(OSI
MiniCPM-o 4.5 is Apache 2.0 and ships its speech decoder with the weights. OpenBMB describes its post-training data and voice recordings but publishes neither, and offers fine-tuning guides rather than its own training recipe.
- https://huggingface.co/api/models?author=openbmb&search=MiniCPM-o&limit=100 recorded 2026-09-27
Every MiniCPM-o checkpoint tagged license:apache-2.0
- https://huggingface.co/api/models/openbmb/MiniCPM-o-4_5?expand[]=downloads&expand[]=cardData&expand[]=gated&expand[]=createdAt&expand[]=lastModified&expand[]=siblings recorded 2026-09-27
Hub metadata for MiniCPM-o-4_5: gated false, four safetensors shards plus a token2wav speech decoder, card license apache-2.0
- https://huggingface.co/openbmb/MiniCPM-o-4_5/raw/main/README.md recorded 2026-09-27
"Built on carefully designed post-training data and professional voice-actor recordings"; "We support fine-tuning with LLaMA-Factory, SWIFT"; "The MiniCPM-o/V model weights and code are open-sourced under the Apache-2.0 license"
Adoption
4 high confidenceHugging Face downloads over the trailing 30 days, summed across OpenBMB's MiniCPM-o 2.6 and 4.5 checkpoints, including its GGUF, AWQ and int4 builds. The search also returns OpenBMB's text-only MiniCPM models, which are excluded.
- https://huggingface.co/api/models?author=openbmb&search=MiniCPM-o&limit=100 recorded 2026-09-27
Six MiniCPM-o checkpoints with 1,151,233 downloads in the trailing 30 days; MiniCPM-o-4_5 721,159 and MiniCPM-o-2_6 344,395
Capability
5 high confidenceMiniCPM-o sees, listens and speaks at the same time, taking continuous video and audio and answering in text and speech, and its Video-MME score matches Qwen3-Omni's. It sits level with Qwen-Omni, with speech in two languages rather than ten.
- https://huggingface.co/openbmb/MiniCPM-o-4_5/raw/main/README.md recorded 2026-09-27
"can process real-time, continuous video and audio input streams simultaneously while generating concurrent text and speech output streams"; OpenCompass 77.6; Video-MME 70.4, Qwen3-Omni-30B-A3B-Instruct 70.5
Verified 2026-09-27