wav2vec 2.0 / HuBERT
MetaMeta's self-supervised speech encoders: wav2vec 2.0 and HuBERT learn speech representations from unlabeled audio and are fine-tuned into English recognizers with a CTC head. They remain a common starting point for speech fine-tuning and forced alignment.
The fairseq repository that holds the training recipes is archived; the checkpoints are still served from the Hub.
Openness
5 high confidence- weights
- open(PyTorch, TensorFlow and safetensors checkpoints on the Hub, ungated, plus fairseq .pt files)
- data
- open(LibriSpeech (CC BY 4.0) and Libri-Light, both public
- code
- open(fairseq pretraining and CTC fine-tuning recipes for wav2vec 2.0 and HuBERT)
- license
- Apache-2.0(the wav2vec 2.0 and HuBERT checkpoints on the Hub
- research-release
- voxpopuli(2021 VoxPopuli checkpoints under CC-BY-NC-4.0, published with a separate dataset paper
The wav2vec 2.0 and HuBERT checkpoints are Apache-2.0, they were trained on public LibriSpeech and Libri-Light audio, and fairseq publishes the pretraining and fine-tuning recipes, so the models can be rebuilt from what is released. Meta's later VoxPopuli checkpoints carry a non-commercial license but belong to a separate research release.
- https://huggingface.co/api/models?author=facebook&search=wav2vec2&limit=1000&expand[]=downloads&expand[]=createdAt&expand[]=tags&expand[]=gated recorded 2026-09-27
Family listing: the wav2vec2 base, large and XLS-R checkpoints tagged "license:apache-2.0"; the *-voxpopuli* checkpoints tagged "license:cc-by-nc-4.0".
- https://huggingface.co/api/models/facebook/wav2vec2-base-960h?expand[]=downloads&expand[]=cardData&expand[]=gated&expand[]=createdAt&expand[]=lastModified&expand[]=tags&expand[]=siblings recorded 2026-09-27
Hub record for wav2vec2-base-960h: "gated":false, tag "license:apache-2.0", and model.safetensors in the file list.
- https://huggingface.co/facebook/wav2vec2-base-960h/raw/main/README.md recorded 2026-09-27
Card: "The base model pretrained and fine-tuned on 960 hours of Librispeech on 16kHz sampled speech audio."
- https://raw.githubusercontent.com/facebookresearch/fairseq/main/examples/hubert/README.md recorded 2026-09-27
HuBERT README: "HuBERT Large (~316M params) | [Libri-Light](https://github.com/facebookresearch/libri-light) 60k hr"; "## Train a new model".
- https://raw.githubusercontent.com/facebookresearch/fairseq/main/examples/wav2vec/README.md recorded 2026-09-27
wav2vec README: "### Train a wav2vec 2.0 base model"; "### Fine-tune a pre-trained model with CTC".
- https://raw.githubusercontent.com/facebookresearch/fairseq/main/README.md recorded 2026-09-27
fairseq README: "fairseq(-py) is MIT-licensed." and "The license applies to the pre-trained models as well."
- https://www.openslr.org/12/ recorded 2026-09-27
OpenSLR LibriSpeech page: "<b>License:</b> CC BY 4.0".
Adoption
4 high confidenceHugging Face downloads of the five declared wav2vec 2.0 and HuBERT checkpoints; the base encoder is pulled mostly as a starting point for fine-tuning.
- https://huggingface.co/api/models?author=facebook&search=hubert&limit=100&expand[]=downloads&expand[]=createdAt&expand[]=tags&expand[]=gated recorded 2026-09-27
Trailing-30-day downloads: hubert-large-ls960-ft 313,174, hubert-base-ls960 243,099.
- https://huggingface.co/api/models?author=facebook&search=wav2vec2&limit=1000&expand[]=downloads&expand[]=createdAt&expand[]=tags&expand[]=gated recorded 2026-09-27
Trailing-30-day downloads: wav2vec2-base 2,161,483, wav2vec2-base-960h 1,421,400, wav2vec2-large-960h-lv60-self 152,118.
Capability
2 low confidenceAn English recognizer from 2020 whose strength today is as a pretrained encoder to fine-tune. Its reported results cover LibriSpeech only, and it has no place on the public boards that rank Whisper, so it sits a step below Whisper.
- https://artificialanalysis.ai/speech-to-text recorded 2026-09-27
The speech-to-text table lists no row for this model.
- https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-results/resolve/d2c5b384deccdb82834f41aeaffcc618c00efa2f/english_short_latest.csv recorded 2026-09-27
The results CSV lists no facebook/wav2vec2 or facebook/hubert row among its 71 models.
- https://huggingface.co/facebook/wav2vec2-base-960h/raw/main/README.md recorded 2026-09-27
Card results table: "| 3.4 | 8.6 |".
- https://huggingface.co/facebook/wav2vec2-large-960h-lv60-self/raw/main/README.md recorded 2026-09-27
Card results table under "clean" and "other": "| 1.9 | 3.9 |".
Verified 2026-09-27