Sinhala NLP
unknownScores
1 product on the map — 1 closed.
SOLD (Sinhala Offensive Language Dataset)
Openness
2 medium confidence- license
- not-clearly-stated-on-card(no license in the Hugging Face card, API metadata or GitHub README
- access
- public
- dataset_card
- present(card describes the annotation scheme and columns)
The posts and labels download freely, but no license appears anywhere the release is published, so reuse terms are unknown. The data also carries Twitter post IDs and text, which the platform's own terms govern.
- https://huggingface.co/api/datasets/sinhala-nlp/SOLD recorded 2026-09-24
gated: false; no license in cardData
- https://huggingface.co/datasets/sinhala-nlp/SOLD/raw/main/README.md recorded 2026-09-24
Card metadata has task_categories and language only; columns include post_id (Twitter ID) and text
- https://raw.githubusercontent.com/Sinhala-NLP/SOLD/master/README.md recorded 2026-09-24
README text with annotation scheme and experiments; no license section
Adoption
1 high confidenceHugging Face downloads summed over SOLD and SemiSOLD. Clones of the experiment code on GitHub are not counted.
- https://huggingface.co/api/datasets/sinhala-nlp/SemiSOLD recorded 2026-09-24
41 downloads in the trailing 30 days for sinhala-nlp/SemiSOLD
- https://huggingface.co/api/datasets/sinhala-nlp/SOLD recorded 2026-09-24
254 downloads in the trailing 30 days for sinhala-nlp/SOLD
Capability
3 medium confidenceSOLD is a documented, hand-labeled set for a single task in one language, offensive language in Sinhala tweets. Nothing outside its paper builds on it, whereas AfroBench, for African languages, gathers 22 datasets into a suite with a public leaderboard.
- https://arxiv.org/abs/2212.00851 recorded 2026-09-24
"SOLD is a manually annotated dataset containing 10,000 posts from Twitter"; "SemiSOLD, a larger dataset containing more than 145,000 Sinhala tweets"