WAXAL
GoogleWAXAL is a speech corpus for Sub-Saharan African languages in two parts: about 1,250 hours of transcribed natural speech for speech recognition in 19 languages, and single-speaker recordings of phonetically balanced scripts for text-to-speech in 17 languages. Partners including Makerere University, the University of Ghana and Digital Umuganda collected the audio. Google Research publishes it.
Openness
5 high confidence- license
- cc-by-sa-4.0(Makerere, Digital Umuganda, Media Trust, Loud and Clear and AIMS Senegal subsets)+cc-by-4.0(University of Ghana subsets)
- access
- public(Hugging Face repository is not gated)
- dataset_card
- present(per-provider tables, fields, splits and speaker-disjoint v2 splits)
Every subset downloads without a gate, and each collecting partner's license allows commercial use with attribution. Most partners add a share-alike condition; the University of Ghana subsets do not.
- https://arxiv.org/abs/2602.02734 recorded 2026-09-24
Abstract: "an openly accessible speech dataset for 24 languages" with about 1,250 hours of ASR speech and 235 hours of TTS recordings.
- https://huggingface.co/api/datasets/google/WaxalNLP recorded 2026-09-24
"gated": false; cardData license ["cc-by-sa-4.0", "cc-by-4.0"].
- https://huggingface.co/datasets/google/WaxalNLP recorded 2026-09-24
Provider tables list each subset's license: Makerere University, Digital Umuganda, Media Trust, Loud and Clear and AIMS Senegal "CC-BY-SA-4.0"; University of Ghana "CC-BY-4.0".
Adoption
3 high confidenceHugging Face downloads of the single WaxalNLP repository, which holds both the ASR and TTS halves. Downloads count file fetches and cannot show how many models were trained on the data.
- https://huggingface.co/api/datasets/google/WaxalNLP recorded 2026-09-24
13562 downloads in the trailing 30 days for google/WaxalNLP
Capability
3 medium confidenceWAXAL pairs transcribed natural speech in nineteen languages with studio text-to-speech recordings in seventeen, all recorded and transcribed by people and documented in a paper, and recognizers such as Ethio-ASR are trained on it, with part of its audio drawn from Afrivoice. Its roughly twelve hundred transcribed hours trail Afrivoice's thousands for many of the same languages and are a sliver of the hundreds of thousands in the largest English speech corpora.
- https://arxiv.org/abs/2602.02734 recorded 2026-09-24
Paper abstract describes collection, annotation and quality control with four African partner organizations.
- https://huggingface.co/datasets/google/WaxalNLP recorded 2026-09-24
Card: ASR "approximately 1,250 hours of transcribed natural speech" in 19 languages; "Models trained or fine-tuned on google/WaxalNLP" lists badrex/Ethio-ASR-multilingual-600M.
Verified 2026-09-24