AI Potluck
Back to Gap Map Model components / Language-specific datasets

Vuk'uzenzele corpus

Data Science for Social Impact, University of Pretoria
open / Overall score: 1.3

The Vuk'uzenzele corpus is text from South Africa's government magazine Vuk'uzenzele, extracted from its PDF editions in all 11 official languages. It ships as a monolingual set with train, test and eval splits and as a sentence-aligned parallel set of more than 60,000 pairs across 55 language combinations, aligned with LASER embeddings. The Data Science for Social Impact group (DSFSI) prepared it.

Openness

5 high confidence
5.0
license
cc-by-4.0(LICENSE.data.md covers the data
access
public(both Hugging Face repositories ungated
dataset_card
present(source, languages, format and disclaimer described

The data is under an attribution-only license and both repositories download without a gate. The source is a government magazine, and the card disclaims accuracy of the PDF extraction rather than restricting use.

Adoption

2 high confidence
2.0

Hugging Face downloads summed over the monolingual and sentence-aligned repositories; the Zenodo archive is not counted. Downloads count file fetches, not users.

Capability

1 medium confidence
1.0

The Vuk'uzenzele corpus covers all eleven official South African languages in government-translated magazine text extracted from PDF editions, is documented in a paper, and its publisher trained PuoBERTa and translation models on it. Its tens of thousands of automatically aligned sentences from a single domain are much smaller than WURA's general web corpus and nowhere near the billions of pairs in the largest parallel collections.

Verified 2026-09-24