AI Potluck
Back to Gap Map Organization

Castorini

lab · Canada

Scores

1 product on the map — 1 open.

WURA

Openness

5 medium confidence
5.0
license
apache-2.0(Hugging Face card metadata
access
public(Hugging Face repository is not gated)
dataset_card
present(short card naming sources (audited mC4, crawled news)

The card applies Apache-2.0 to the corpus, which downloads without a gate. The text is crawled from the web and news sites, and the card does not address the original publishers' terms.

Adoption

2 high confidence
2.0

Hugging Face downloads of the single WURA repository. Downloads count file fetches and cannot show how many models were pretrained on it.

Capability

3 medium confidence
3.0

WURA is a documented web corpus for sixteen African languages whose builders audited Google's multilingual web crawl by hand and added verified news, and the second-generation AfriTeVa models are pretrained on it, though it is crawled text rather than new writing and several languages have only tens of thousands of documents. At roughly ten billion tokens, estimated from its size on disk, it is a small fraction of English web corpora such as FineWeb, which reach trillions.