WhisperX
m-bainA transcription pipeline around Whisper that adds accurate word-level timestamps through wav2vec 2.0 forced alignment, speaker labels through pyannote.audio diarization, and voice-activity-based batching on the faster-whisper backend, for up to 70 times real-time transcription.
It composes faster-whisper, wav2vec 2.0 alignment models and pyannote.audio, all separate entries here; the diarization model it downloads sits behind a Hugging Face user agreement.
Openness
5 medium confidence- license
- BSD-2-Clause(OSI)
- source
- public
- core features withheld
- no — an individual researcher's project from Oxford's Visual Geometry Group, with no paid tier
WhisperX is BSD-2-Clause and the published repository is the whole pipeline. It is maintained by its author out of academic research, and nothing is sold alongside it.
- https://raw.githubusercontent.com/m-bain/whisperX/main/LICENSE recorded 2026-09-27
LICENSE body: "BSD 2-Clause License" / "Copyright (c) 2024, Max Bain".
- https://raw.githubusercontent.com/m-bain/whisperX/main/README.md recorded 2026-09-27
README: "This work, and my PhD, is supported by the [VGG (Visual Geometry Group)](https://www.robots.ox.ac.uk/~vgg/) and the University of Oxford."
Adoption
3 high confidencePyPI downloads of the whisperx package.
- https://pypistats.org/api/packages/whisperx/recent recorded 2026-09-27
"last_month":328332 for whisperx.
Capability
3 medium confidenceThe usual way to get speaker-labeled, word-timed Whisper transcripts, but it runs one task by composing other engines and trains nothing, two steps below NeMo.
- https://raw.githubusercontent.com/m-bain/whisperX/main/README.md recorded 2026-09-27
README: "This repository provides fast automatic speech recognition (70x realtime with large-v2) with word-level timestamps and speaker diarization."
Verified 2026-09-27