AI Potluck
Back to Gap Map Model components / Speech & audio

wav2vec 2.0 / HuBERT

Meta
open source / Overall score: 2.8

Meta's self-supervised speech encoders: wav2vec 2.0 and HuBERT learn speech representations from unlabeled audio and are fine-tuned into English recognizers with a CTC head. They remain a common starting point for speech fine-tuning and forced alignment.

The fairseq repository that holds the training recipes is archived; the checkpoints are still served from the Hub.

Openness

5 high confidence
5.0
weights
open(PyTorch, TensorFlow and safetensors checkpoints on the Hub, ungated, plus fairseq .pt files)
data
open(LibriSpeech (CC BY 4.0) and Libri-Light, both public
code
open(fairseq pretraining and CTC fine-tuning recipes for wav2vec 2.0 and HuBERT)
license
Apache-2.0(the wav2vec 2.0 and HuBERT checkpoints on the Hub
research-release
voxpopuli(2021 VoxPopuli checkpoints under CC-BY-NC-4.0, published with a separate dataset paper

The wav2vec 2.0 and HuBERT checkpoints are Apache-2.0, they were trained on public LibriSpeech and Libri-Light audio, and fairseq publishes the pretraining and fine-tuning recipes, so the models can be rebuilt from what is released. Meta's later VoxPopuli checkpoints carry a non-commercial license but belong to a separate research release.

Adoption

4 high confidence
4.0

Hugging Face downloads of the five declared wav2vec 2.0 and HuBERT checkpoints; the base encoder is pulled mostly as a starting point for fine-tuning.

Capability

2 low confidence
2.0

An English recognizer from 2020 whose strength today is as a pretrained encoder to fine-tune. Its reported results cover LibriSpeech only, and it has no place on the public boards that rank Whisper, so it sits a step below Whisper.

Verified 2026-09-27