AI Potluck
Back to Gap Map Model components / Language-specific datasets

AfriDocMT

Masakhane
restricted / Overall score: 1.0

AfriDocMT is a translation dataset of whole documents in which 334 health and 271 information technology news documents were translated by people from English into Amharic, Hausa, Swahili, Yoruba and isiZulu. It ships in both document and sentence form, with train, dev and test splits for each domain. Its builders were funded by the Lacuna Fund, and Masakhane hosts the release.

Openness

2 high confidence
2.0
license
cc-by-nc-sa-3.0(Health documents)+cc-by-nc-4.0(Tech documents)
access
public(Hugging Face repository is not gated)
dataset_card
present(short card with layout, languages, domains and licenses

The data downloads without a gate, but both licenses bar commercial use: share-alike non-commercial terms for the health documents and non-commercial terms for the technology ones.

Adoption

1 high confidence
1.0

Hugging Face downloads of the single repository. Downloads count file fetches and cannot separate training use from evaluation.

Capability

1 medium confidence
1.0

AfriDocMT gives five African languages whole health and IT news documents translated by people from English, documented in a paper, though no model or benchmark beyond its own experiments is built on it. At a few hundred documents it is minute beside the billions of pairs in the largest parallel collections, and far smaller than the IIT Bombay English-Hindi corpus for one pair.

  • https://arxiv.org/abs/2501.06374 recorded 2026-09-24

    Abstract: "334 health and 271 information technology news documents, all human-translated from English"; NMT and LLM benchmark experiments.

Verified 2026-09-24