NCHLT named-entity corpora
Centre for Text Technology, North-West UniversityThe NCHLT text corpora are foundational text resources for South African languages: monolingual corpora with subsets annotated for tokenization, morphology and part of speech, used to build tokenizers, lemmatizers and taggers. The Hugging Face entry loads the named-entity annotated corpora for nine languages, from Afrikaans to isiZulu, from the SADiLaR repository. A North-West University curator maintains the entry.
Scored as the Hugging Face loader for the named-entity annotated corpora; the other NCHLT corpora on SADiLaR are not separately listed.
Openness
4 medium confidence- license
- cc-by-2.5(Hugging Face card metadata
- access
- public(ungated loader
- dataset_card
- partial(paper abstract only
The loader is ungated and pulls openly downloadable archives from SADiLaR under an attribution license. The card itself documents almost nothing beyond the paper abstract.
- https://huggingface.co/api/datasets/nwu-ctext/nchlt recorded 2026-09-24
"gated": false; repository files .gitattributes, README.md, nchlt.py.
- https://huggingface.co/datasets/nwu-ctext/nchlt recorded 2026-09-24
Card sections read "[More Information Needed]"; curator "Martin.Puttkammer@nwu.ac.za".
- https://huggingface.co/datasets/nwu-ctext/nchlt/raw/main/nchlt.py recorded 2026-09-24
Configs download e.g. "https://repo.sadilar.org/bitstream/handle/20.500.12185/299/nchlt_afrikaans_named_entity_annotated_corpus.zip".
- https://huggingface.co/datasets/nwu-ctext/nchlt/raw/main/README.md recorded 2026-09-24
Front matter "license: cc-by-2.5", nine languages, "named-entity-recognition".
Adoption
1 high confidenceHugging Face downloads of the loader repository only; direct downloads from SADiLaR are not counted. Downloads count loader fetches, not users.
- https://huggingface.co/api/datasets/nwu-ctext/nchlt recorded 2026-09-24
166 downloads in the trailing 30 days for nwu-ctext/nchlt
Capability
3 medium confidenceThe NCHLT corpora are an established resource for South African languages, described in an LREC paper, and PuoBERTa lists them among its training data. The Hugging Face entry exposes only the named-entity subset through a loader with an empty card, so it offers far less text than WURA, a documented African web corpus that trained the AfriTeVa V2 models.
- https://huggingface.co/datasets/nwu-ctext/nchlt recorded 2026-09-24
Models trained or fine-tuned list includes dsfsi/PuoBERTa.
- https://huggingface.co/datasets/nwu-ctext/nchlt/raw/main/README.md recorded 2026-09-24
config af train num_examples 8961; NER tag set B-PERS, B-ORG, B-LOC, B-MISC.
Verified 2026-09-24