AI Potluck
Back to Gap Map Model components / Language-specific datasets

NCHLT named-entity corpora

Centre for Text Technology, North-West University
open / Overall score: 2.4

The NCHLT text corpora are foundational text resources for South African languages: monolingual corpora with subsets annotated for tokenization, morphology and part of speech, used to build tokenizers, lemmatizers and taggers. The Hugging Face entry loads the named-entity annotated corpora for nine languages, from Afrikaans to isiZulu, from the SADiLaR repository. A North-West University curator maintains the entry.

Scored as the Hugging Face loader for the named-entity annotated corpora; the other NCHLT corpora on SADiLaR are not separately listed.

Openness

4 medium confidence
4.0
license
cc-by-2.5(Hugging Face card metadata
access
public(ungated loader
dataset_card
partial(paper abstract only

The loader is ungated and pulls openly downloadable archives from SADiLaR under an attribution license. The card itself documents almost nothing beyond the paper abstract.

Adoption

1 high confidence
1.0

Hugging Face downloads of the loader repository only; direct downloads from SADiLaR are not counted. Downloads count loader fetches, not users.

Capability

3 medium confidence
3.0

The NCHLT corpora are an established resource for South African languages, described in an LREC paper, and PuoBERTa lists them among its training data. The Hugging Face entry exposes only the named-entity subset through a loader with an empty card, so it offers far less text than WURA, a documented African web corpus that trained the AfriTeVa V2 models.

Verified 2026-09-24