AfriHate
AfriHateAfriHate is a hate speech and abusive language dataset of tweets in 15 African languages, including Algerian and Moroccan Arabic, Amharic, Hausa, Nigerian Pidgin, Swahili, Yoruba and isiZulu. Native speakers familiar with the local culture labeled each tweet as hate, abusive or normal. The accompanying paper reports classification baselines from fine-tuned encoders such as AfriBERTa to prompted LLMs such as Llama 3.1 and Gemma 2.
Openness
3 medium confidence- license
- apache-2.0(Hugging Face card metadata
- access
- auto(automatic Hugging Face gate)
- dataset_card
- present(languages, labels, fields and baselines described
The card metadata applies Apache-2.0, and the data opens after an automatic click-through on Hugging Face. The GitHub repository carries no license file of its own.
- https://huggingface.co/api/datasets/afrihate/afrihate recorded 2026-09-24
"gated": "auto"; cardData license "apache-2.0".
- https://huggingface.co/datasets/afrihate/afrihate recorded 2026-09-24
"Each example in the dataset is a tweet annotated by native speakers"; labels "Hate, Abusive, or Normal"; 15 languages listed.
- https://raw.githubusercontent.com/afrihate/afrihate/main/README.md recorded 2026-09-24
README points to the paper and the Hugging Face dataset; the license badge is commented out.
Adoption
1 high confidenceHugging Face downloads of the single AfriHate repository, behind an automatic gate. Downloads count file fetches, not moderation systems built on it.
- https://huggingface.co/api/datasets/afrihate/afrihate recorded 2026-09-24
134 downloads in the trailing 30 days for afrihate/afrihate
Capability
3 medium confidenceAfriHate is a documented, native-speaker-labeled dataset for one task, hate and abuse detection, across 15 languages. Beyond its own baselines, only a small community classifier is listed as trained on it, and it covers a single task where AfroBench spans 15 tasks in 64 African languages.
- https://arxiv.org/abs/2501.08284 recorded 2026-09-24
Abstract: "a multilingual collection of hate speech and abusive language datasets in 15 African languages. Each instance in AfriHate is annotated by native speakers".
- https://arxiv.org/html/2501.08284 recorded 2026-09-24
"each tweet was annotated by 3 annotators" for languages with pre-annotation; tweets from 2012 to 2023 via the Academic API.
Verified 2026-09-24