Mozilla Foundation
foundation · United StatesScores
1 product on the map — 1 open-ish.
Openness
3 medium confidence- license
- cc0-1.0(Per-language dataset pages on Mozilla Data Collective list CC0-1.0
- access
- terms-acceptance+contact-info(Downloads go through a Mozilla Data Collective account and API key after agreeing to each dataset's conditions, which forbid re-hosting and speaker identification.)
- dataset_card
- present(Per-language datasheets on MDC plus release documentation in cv-dataset.)
The recordings are dedicated to the public domain, so any use is allowed. Getting them takes a Mozilla Data Collective account and agreement to conditions that forbid re-sharing the files or trying to identify speakers.
- https://mozilladatacollective.com/datasets/cmu5o9t9r00xfmh078ono52dr recorded 2026-09-24
Per-language page: "License: CC0-1.0"; "It is forbidden to attempt to determine the identity of speakers... It is forbidden to re-host or re-share this dataset."
- https://mozilladatacollective.com/organization/cmfh0j9o10006ns07jq45h7xk recorded 2026-09-24
Mozilla Data Collective lists every Common Voice release by language, each with "CC0-1.0" as the dataset license.
- https://raw.githubusercontent.com/common-voice/cv-dataset/main/LICENSE recorded 2026-09-24
Mozilla Public License Version 2.0 for the cv-dataset repository.
- https://raw.githubusercontent.com/common-voice/cv-dataset/main/README.md recorded 2026-09-24
"Please visit the Mozilla Data Collective Common Voice section to download the latest datasets." Documents dataset types, fields and the data pipeline.
- https://raw.githubusercontent.com/Mozilla-Data-Collective/datacollective-python/main/README.md recorded 2026-09-24
"Before trying to access any dataset, make sure you have thoroughly read and agreed to the specific dataset's conditions & licensing terms." "Get your API key from the Mozilla Data Collective dashboard".
Adoption
not assessedCommon Voice is distributed only through Mozilla Data Collective, and its dataset page there shows no download figure. The Hugging Face repositories are stubs pointing to it, so their counts would not measure use of the corpus.
- https://mozilladatacollective.com/datasets/cmu5o9t9r00xfmh078ono52dr recorded 2026-09-24
The Common Voice dataset page offers download sessions and API access and carries no download count.
Capability
4 medium confidenceCommon Voice gives nearly three hundred languages public-domain read speech recorded and checked by volunteers and documented language by language, and Mozilla's DeepSpeech experiments and the comparison recognizers in Meta's MMS work are trained on it, doing better on FLEURS than models trained on Bible readings. Its tens of thousands of validated hours make it the broadest open speech collection, though the biggest English speech corpora run to hundreds of thousands of hours.
- https://arxiv.org/abs/1912.06670 recorded 2026-09-24
"the largest audio corpus in the public domain for speech recognition, both in terms of number of hours and number of languages"; DeepSpeech experiments.
- https://arxiv.org/pdf/2305.13516 recorded 2026-09-24
"models trained on CommonVoice perform better on 18 languages of FLEURS (average CER 9.3 vs. 12.2)".
- https://raw.githubusercontent.com/common-voice/cv-dataset/main/README.md recorded 2026-09-24
Scripted Speech: 27 releases, latest v27.0, 295 languages; hours chart ends at 42,593 total and 29,295 validated.