AI Potluck
Back to Gap Map Model components / Speech & audio

WhisperX

m-bain
open source / Overall score: 3.0

A transcription pipeline around Whisper that adds accurate word-level timestamps through wav2vec 2.0 forced alignment, speaker labels through pyannote.audio diarization, and voice-activity-based batching on the faster-whisper backend, for up to 70 times real-time transcription.

It composes faster-whisper, wav2vec 2.0 alignment models and pyannote.audio, all separate entries here; the diarization model it downloads sits behind a Hugging Face user agreement.

Openness

5 medium confidence
5.0
license
BSD-2-Clause(OSI)
source
public
core features withheld
no — an individual researcher's project from Oxford's Visual Geometry Group, with no paid tier

WhisperX is BSD-2-Clause and the published repository is the whole pipeline. It is maintained by its author out of academic research, and nothing is sold alongside it.

Adoption

3 high confidence
3.0

PyPI downloads of the whisperx package.

Capability

3 medium confidence
3.0

The usual way to get speaker-labeled, word-timed Whisper transcripts, but it runs one task by composing other engines and trains nothing, two steps below NeMo.

Verified 2026-09-27