AI Potluck
Back to Gap Map Model components / Language-specific datasets

SOLD (Sinhala Offensive Language Dataset)

Sinhala NLP
restricted / Overall score: 2.4

SOLD is a Sinhala offensive language dataset of 10,000 Twitter posts labeled by hand as offensive or not, with rationales marking the words behind each offensive label. Its companion SemiSOLD adds more than 145,000 Sinhala tweets scored by classifiers trained on SOLD. The Sinhala-NLP group released both with an experiment codebase.

Openness

2 medium confidence
2.0
license
not-clearly-stated-on-card(no license in the Hugging Face card, API metadata or GitHub README
access
public
dataset_card
present(card describes the annotation scheme and columns)

The posts and labels download freely, but no license appears anywhere the release is published, so reuse terms are unknown. The data also carries Twitter post IDs and text, which the platform's own terms govern.

Adoption

1 high confidence
1.0

Hugging Face downloads summed over SOLD and SemiSOLD. Clones of the experiment code on GitHub are not counted.

Capability

3 medium confidence
3.0

SOLD is a documented, hand-labeled set for a single task in one language, offensive language in Sinhala tweets. Nothing outside its paper builds on it, whereas AfroBench, for African languages, gathers 22 datasets into a suite with a public leaderboard.

  • https://arxiv.org/abs/2212.00851 recorded 2026-09-24

    "SOLD is a manually annotated dataset containing 10,000 posts from Twitter"; "SemiSOLD, a larger dataset containing more than 145,000 Sinhala tweets"

Verified 2026-09-24