AI Potluck
Back to Gap Map Model components / Language-specific datasets

WURA

Castorini
open / Overall score: 2.7

WURA is a pretraining corpus of whole web and news documents in 16 African languages plus Arabic, English, French and Portuguese. Its builders audited mC4 for thirteen African languages to strip low-quality text and crawled additional verified news sources, keeping headlines and source URLs with each document. Castorini released it alongside the AfriTeVa V2 models it was built to train.

Openness

5 medium confidence
5.0
license
apache-2.0(Hugging Face card metadata
access
public(Hugging Face repository is not gated)
dataset_card
present(short card naming sources (audited mC4, crawled news)

The card applies Apache-2.0 to the corpus, which downloads without a gate. The text is crawled from the web and news sites, and the card does not address the original publishers' terms.

Adoption

2 high confidence
2.0

Hugging Face downloads of the single WURA repository. Downloads count file fetches and cannot show how many models were pretrained on it.

Capability

3 medium confidence
3.0

WURA is a documented web corpus for sixteen African languages whose builders audited Google's multilingual web crawl by hand and added verified news, and the second-generation AfriTeVa models are pretrained on it, though it is crawled text rather than new writing and several languages have only tens of thousands of documents. At roughly ten billion tokens, estimated from its size on disk, it is a small fraction of English web corpora such as FineWeb, which reach trillions.

Verified 2026-09-24