SIB-200
David AdelaniSIB-200 is a topic classification benchmark in 205 languages and dialects, built by labeling the English FLORES-200 sentences with seven topics such as science, travel and politics and carrying the labels across to every aligned translation. Train, validation and test splits exist for every language. It was created by David Adelani and colleagues, with annotators recruited through Masakhane.
Openness
5 medium confidence- license
- cc-by-sa-4.0(Hugging Face license tag
- access
- public
- dataset_card
- present(Describes composition and splits, but several sections are unfilled template text that wrongly calls the source news.)
The data downloads from Hugging Face without a gate under a share-alike license tag. The card's own wording of the license is loose and the GitHub license covers the code, so the tag is the clearest statement.
- https://huggingface.co/api/datasets/Davlan/sib200 recorded 2026-09-24
gated: false; license tag cc-by-sa-4.0.
- https://huggingface.co/datasets/Davlan/sib200/raw/main/README.md recorded 2026-09-24
"The licensing status of the data is CC 4.0 Commercial"; "Annotators were recruited from Masakhane".
- https://raw.githubusercontent.com/dadelani/sib-200/main/LICENSE recorded 2026-09-24
Apache License, Version 2.0.
Adoption
3 high confidenceHugging Face downloads over the trailing month for the SIB-200 repository. Rebuilding it from FLORES with the GitHub script is not counted.
- https://huggingface.co/api/datasets/Davlan/sib200 recorded 2026-09-24
22383 downloads in the trailing 30 days for Davlan/sib200
Capability
3 medium confidenceSIB-200 is documented in its paper and covers more languages than Belebele. It is a single topic-labeling task over FLORES sentences, and no named model or leaderboard beyond its own baselines was found to use it.
- https://arxiv.org/abs/2309.07445 recorded 2026-09-24
"We annotated the English portion of the dataset and extended the ... annotation to the remaining 203 languages covered in the corpus".
- https://huggingface.co/datasets/Davlan/sib200/raw/main/README.md recorded 2026-09-24
"SIB-200 is the largest publicly available topic classification dataset based on Flores-200 covering 205 languages and dialects"; English split 701/99/204.
Verified 2026-09-24