MOSS-Transcribe-Diarize
OpenMOSSOpenMOSS's end-to-end model for long multi-speaker recordings: in one pass over up to 90 minutes of audio it transcribes, labels who spoke when, and adds timestamps and acoustic events, in more than 50 languages. It pairs a Whisper-style encoder with a small Qwen3-style decoder.
Openness
3 medium confidence- weights
- open(0.9B safetensors checkpoint on the Hub, ungated)
- data
- closed(trained on in-the-wild recordings that are not listed or released)
- code
- partial(inference, serving and a minimal fine-tuning script)
- license
- Apache-2.0(code and the 0.9B weights)
- api-only-tier
- MOSS-Transcribe-Diarize Pro(a stronger model offered only through MOSI's online playground
The 0.9B model and its code are Apache-2.0 and download freely, but the training recordings are neither listed nor released and only a minimal fine-tuning script is published. A stronger Pro model is offered only in MOSI's hosted playground.
- https://arxiv.org/abs/2601.01554 recorded 2026-09-27
Abstract: "Trained on extensive real wild data and equipped with a 128k context window for up to 90-minute inputs".
- https://huggingface.co/api/models/OpenMOSS-Team/MOSS-Transcribe-Diarize?expand[]=downloads&expand[]=cardData&expand[]=gated&expand[]=createdAt&expand[]=lastModified&expand[]=tags&expand[]=siblings recorded 2026-09-27
Hub record: "gated":false and model-00000-of-00001.safetensors in the file list.
- https://huggingface.co/OpenMOSS-Team/MOSS-Transcribe-Diarize/raw/main/README.md recorded 2026-09-27
Card: "MOSS-Transcribe-Diarize 0.9B is licensed under the Apache License 2.0."
- https://raw.githubusercontent.com/OpenMOSS/MOSS-Transcribe-Diarize/main/FINETUNING.md recorded 2026-09-27
FINETUNING.md: "`finetune.py` provides a minimal fine-tuning workflow built on the Hugging Face Transformers `Trainer`."
- https://raw.githubusercontent.com/OpenMOSS/MOSS-Transcribe-Diarize/main/README.md recorded 2026-09-27
README: "* 2026-07-09: Open-sourced MOSS-Transcribe-Diarize 0.9B."; "[MOSS-Transcribe-Diarize Pro](https://platform.mosi.cn/app/playground) is a stronger model with higher overall performance and is available through the online playground."
Adoption
3 medium confidenceHugging Face downloads of the 0.9B checkpoint within about three months of its open release.
- https://huggingface.co/api/models?author=OpenMOSS-Team&search=MOSS-Transcribe&limit=1000&expand[]=downloads&expand[]=createdAt&expand[]=tags&expand[]=gated recorded 2026-09-27
Trailing-30-day downloads: OpenMOSS-Team/MOSS-Transcribe-Diarize 165,491.
Capability
4 medium confidenceClose to the best open recognizers on English accuracy while also diarizing long recordings, a step below Qwen3-ASR.
- https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-results/resolve/d2c5b384deccdb82834f41aeaffcc618c00efa2f/english_short_latest.csv recorded 2026-09-27
Results CSV row "OpenMOSS-Team/MOSS-Transcribe-Diarize,4.63625".
- https://raw.githubusercontent.com/OpenMOSS/MOSS-Transcribe-Diarize/main/README.md recorded 2026-09-27
README: "an open-source SOTA end-to-end audio understanding model for long-form multi-speaker transcription, diarization, timestamps, and acoustic event awareness".
Verified 2026-09-27