AI Potluck
Back to Gap Map Model components / Speech & audio

MOSS-Transcribe-Diarize

OpenMOSS
open weights / Overall score: 3.6

OpenMOSS's end-to-end model for long multi-speaker recordings: in one pass over up to 90 minutes of audio it transcribes, labels who spoke when, and adds timestamps and acoustic events, in more than 50 languages. It pairs a Whisper-style encoder with a small Qwen3-style decoder.

Openness

3 medium confidence
3.0
weights
open(0.9B safetensors checkpoint on the Hub, ungated)
data
closed(trained on in-the-wild recordings that are not listed or released)
code
partial(inference, serving and a minimal fine-tuning script)
license
Apache-2.0(code and the 0.9B weights)
api-only-tier
MOSS-Transcribe-Diarize Pro(a stronger model offered only through MOSI's online playground

The 0.9B model and its code are Apache-2.0 and download freely, but the training recordings are neither listed nor released and only a minimal fine-tuning script is published. A stronger Pro model is offered only in MOSI's hosted playground.

Adoption

3 medium confidence
3.0

Hugging Face downloads of the 0.9B checkpoint within about three months of its open release.

Capability

4 medium confidence
4.0

Close to the best open recognizers on English accuracy while also diarizing long recordings, a step below Qwen3-ASR.

Verified 2026-09-27