AI Potluck
Back to Gap Map Organization

Sinhala NLP

unknown

Scores

1 product on the map — 1 closed.

SOLD (Sinhala Offensive Language Dataset)

Openness

2 medium confidence
2.0
license
not-clearly-stated-on-card(no license in the Hugging Face card, API metadata or GitHub README
access
public
dataset_card
present(card describes the annotation scheme and columns)

The posts and labels download freely, but no license appears anywhere the release is published, so reuse terms are unknown. The data also carries Twitter post IDs and text, which the platform's own terms govern.

Adoption

1 high confidence
1.0

Hugging Face downloads summed over SOLD and SemiSOLD. Clones of the experiment code on GitHub are not counted.

Capability

3 medium confidence
3.0

SOLD is a documented, hand-labeled set for a single task in one language, offensive language in Sinhala tweets. Nothing outside its paper builds on it, whereas AfroBench, for African languages, gathers 22 datasets into a suite with a public leaderboard.

  • https://arxiv.org/abs/2212.00851 recorded 2026-09-24

    "SOLD is a manually annotated dataset containing 10,000 posts from Twitter"; "SemiSOLD, a larger dataset containing more than 145,000 Sinhala tweets"