AI Potluck
Back to Gap Map Model components / Language-specific datasets

WAXAL

Google
open / Overall score: 3.0

WAXAL is a speech corpus for Sub-Saharan African languages in two parts: about 1,250 hours of transcribed natural speech for speech recognition in 19 languages, and single-speaker recordings of phonetically balanced scripts for text-to-speech in 17 languages. Partners including Makerere University, the University of Ghana and Digital Umuganda collected the audio. Google Research publishes it.

Openness

5 high confidence
5.0
license
cc-by-sa-4.0(Makerere, Digital Umuganda, Media Trust, Loud and Clear and AIMS Senegal subsets)+cc-by-4.0(University of Ghana subsets)
access
public(Hugging Face repository is not gated)
dataset_card
present(per-provider tables, fields, splits and speaker-disjoint v2 splits)

Every subset downloads without a gate, and each collecting partner's license allows commercial use with attribution. Most partners add a share-alike condition; the University of Ghana subsets do not.

Adoption

3 high confidence
3.0

Hugging Face downloads of the single WaxalNLP repository, which holds both the ASR and TTS halves. Downloads count file fetches and cannot show how many models were trained on the data.

Capability

3 medium confidence
3.0

WAXAL pairs transcribed natural speech in nineteen languages with studio text-to-speech recordings in seventeen, all recorded and transcribed by people and documented in a paper, and recognizers such as Ethio-ASR are trained on it, with part of its audio drawn from Afrivoice. Its roughly twelve hundred transcribed hours trail Afrivoice's thousands for many of the same languages and are a sliver of the hundreds of thousands in the largest English speech corpora.

  • https://arxiv.org/abs/2602.02734 recorded 2026-09-24

    Paper abstract describes collection, annotation and quality control with four African partner organizations.

  • https://huggingface.co/datasets/google/WaxalNLP recorded 2026-09-24

    Card: ASR "approximately 1,250 hours of transcribed natural speech" in 19 languages; "Models trained or fine-tuned on google/WaxalNLP" lists badrex/Ethio-ASR-multilingual-600M.

Verified 2026-09-24