Vuk'uzenzele corpus
Data Science for Social Impact, University of PretoriaThe Vuk'uzenzele corpus is text from South Africa's government magazine Vuk'uzenzele, extracted from its PDF editions in all 11 official languages. It ships as a monolingual set with train, test and eval splits and as a sentence-aligned parallel set of more than 60,000 pairs across 55 language combinations, aligned with LASER embeddings. The Data Science for Social Impact group (DSFSI) prepared it.
Openness
5 high confidence- license
- cc-by-4.0(LICENSE.data.md covers the data
- access
- public(both Hugging Face repositories ungated
- dataset_card
- present(source, languages, format and disclaimer described
The data is under an attribution-only license and both repositories download without a gate. The source is a government magazine, and the card disclaims accuracy of the PDF extraction rather than restricting use.
- https://huggingface.co/api/datasets/dsfsi/vukuzenzele-monolingual recorded 2026-09-24
"gated": false; cardData license "cc-by-4.0".
- https://raw.githubusercontent.com/dsfsi/vukuzenzele-nlp/master/LICENSE.data.md recorded 2026-09-24
"Attribution 4.0 International" Creative Commons license text for the data.
- https://raw.githubusercontent.com/dsfsi/vukuzenzele-nlp/master/LICENSE.md recorded 2026-09-24
"MIT License Copyright (c) 2022 Data Science for Social Impact Research Group @ University of Pretoria" for the software.
- https://raw.githubusercontent.com/dsfsi/vukuzenzele-nlp/master/README.md recorded 2026-09-24
Lists the monolingual and sentence-aligned Hugging Face datasets and a Zenodo archive; "Total aligned pairs: 60,000+ sentence pairs across 55 language pair combinations".
Adoption
2 high confidenceHugging Face downloads summed over the monolingual and sentence-aligned repositories; the Zenodo archive is not counted. Downloads count file fetches, not users.
- https://huggingface.co/api/datasets/dsfsi/vukuzenzele-monolingual recorded 2026-09-24
3277 downloads in the trailing 30 days for dsfsi/vukuzenzele-monolingual
- https://huggingface.co/api/datasets/dsfsi/vukuzenzele-sentence-aligned recorded 2026-09-24
2504 downloads in the trailing 30 days for dsfsi/vukuzenzele-sentence-aligned
Capability
1 medium confidenceThe Vuk'uzenzele corpus covers all eleven official South African languages in government-translated magazine text extracted from PDF editions, is documented in a paper, and its publisher trained PuoBERTa and translation models on it. Its tens of thousands of automatically aligned sentences from a single domain are much smaller than WURA's general web corpus and nowhere near the billions of pairs in the largest parallel collections.
- https://huggingface.co/datasets/dsfsi/vukuzenzele-monolingual recorded 2026-09-24
Models trained or fine-tuned list includes dsfsi/PuoBERTa and dsfsi/za-lid-bert.
- https://raw.githubusercontent.com/dsfsi/vukuzenzele-nlp/master/README.md recorded 2026-09-24
"Processed & aligned data: 84 editions"; "Languages: 11 South African official languages".
Verified 2026-09-24