AI Potluck
Back to Gap Map Organization

Centre for Text Technology, North-West University

lab · South Africa

Scores

1 product on the map — 1 open.

NCHLT named-entity corpora

Openness

4 medium confidence
4.0
license
cc-by-2.5(Hugging Face card metadata
access
public(ungated loader
dataset_card
partial(paper abstract only

The loader is ungated and pulls openly downloadable archives from SADiLaR under an attribution license. The card itself documents almost nothing beyond the paper abstract.

Adoption

1 high confidence
1.0

Hugging Face downloads of the loader repository only; direct downloads from SADiLaR are not counted. Downloads count loader fetches, not users.

Capability

3 medium confidence
3.0

The NCHLT corpora are an established resource for South African languages, described in an LREC paper, and PuoBERTa lists them among its training data. The Hugging Face entry exposes only the named-entity subset through a loader with an empty card, so it offers far less text than WURA, a documented African web corpus that trained the AfriTeVa V2 models.