Castorini
lab · CanadaScores
1 product on the map — 1 open.
Openness
5 medium confidence- license
- apache-2.0(Hugging Face card metadata
- access
- public(Hugging Face repository is not gated)
- dataset_card
- present(short card naming sources (audited mC4, crawled news)
The card applies Apache-2.0 to the corpus, which downloads without a gate. The text is crawled from the web and news sites, and the card does not address the original publishers' terms.
- https://huggingface.co/api/datasets/castorini/wura recorded 2026-09-24
"gated": false.
- https://huggingface.co/datasets/castorini/wura recorded 2026-09-24
"This dataset was created by auditing mC4 and crawling additional verified news sources. It was first used to train AfriTeVa V2."
- https://huggingface.co/datasets/castorini/wura/raw/main/README.md recorded 2026-09-24
Front matter "license: apache-2.0"; 20 language configs with headline, content, category and url fields.
Adoption
2 high confidenceHugging Face downloads of the single WURA repository. Downloads count file fetches and cannot show how many models were pretrained on it.
- https://huggingface.co/api/datasets/castorini/wura recorded 2026-09-24
1532 downloads in the trailing 30 days for castorini/wura
Capability
3 medium confidenceWURA is a documented web corpus for sixteen African languages whose builders audited Google's multilingual web crawl by hand and added verified news, and the second-generation AfriTeVa models are pretrained on it, though it is crawled text rather than new writing and several languages have only tens of thousands of documents. At roughly ten billion tokens, estimated from its size on disk, it is a small fraction of English web corpora such as FineWeb, which reach trillions.
- https://huggingface.co/datasets/castorini/wura/raw/main/README.md recorded 2026-09-24
Per-language example counts, e.g. orm train 20,169 and ibo train 51,386 documents.
- https://raw.githubusercontent.com/castorini/AfriTeVa-keji/asiwaju/README.md recorded 2026-09-24
"AfriTeVa-Keji was pretrained on the Wúrà dataset"; releases AfriTeVa V2 Base (428M) and Large (1B).