AfriDocMT
MasakhaneAfriDocMT is a translation dataset of whole documents in which 334 health and 271 information technology news documents were translated by people from English into Amharic, Hausa, Swahili, Yoruba and isiZulu. It ships in both document and sentence form, with train, dev and test splits for each domain. Its builders were funded by the Lacuna Fund, and Masakhane hosts the release.
Openness
2 high confidence- license
- cc-by-nc-sa-3.0(Health documents)+cc-by-nc-4.0(Tech documents)
- access
- public(Hugging Face repository is not gated)
- dataset_card
- present(short card with layout, languages, domains and licenses
The data downloads without a gate, but both licenses bar commercial use: share-alike non-commercial terms for the health documents and non-commercial terms for the technology ones.
- https://huggingface.co/api/datasets/masakhane/AfriDocMT recorded 2026-09-24
"gated": false.
- https://huggingface.co/datasets/masakhane/AfriDocMT/raw/main/README.md recorded 2026-09-24
"License: CC BY-NC-SA 3.0 for Health, and CC BY-NC 4.0 for Tech"; "all human-translated from English".
Adoption
1 high confidenceHugging Face downloads of the single repository. Downloads count file fetches and cannot separate training use from evaluation.
- https://huggingface.co/api/datasets/masakhane/AfriDocMT recorded 2026-09-24
428 downloads in the trailing 30 days for masakhane/AfriDocMT
Capability
1 medium confidenceAfriDocMT gives five African languages whole health and IT news documents translated by people from English, documented in a paper, though no model or benchmark beyond its own experiments is built on it. At a few hundred documents it is minute beside the billions of pairs in the largest parallel collections, and far smaller than the IIT Bombay English-Hindi corpus for one pair.
- https://arxiv.org/abs/2501.06374 recorded 2026-09-24
Abstract: "334 health and 271 information technology news documents, all human-translated from English"; NMT and LLM benchmark experiments.
Verified 2026-09-24