SEACrowd
SEACrowdSEACrowd collects and standardizes datasets for Southeast Asian languages behind one Python library, `seacrowd`, whose dataloaders return text, speech and vision-language corpora in shared schemas with license and citation metadata attached. The hub covers nearly 1,000 languages of the region across three modalities and ships benchmark bundles such as SEACrowd-VL. It is a grassroots research community effort modeled on NusaCrowd.
Scored as the hub and its library. The datasets it loads keep their own publishers and licenses, and the Hugging Face mirrors under the SEACrowd account are loaders for those datasets rather than new releases.
Openness
4 high confidence- license
- mixed-per-subset(The loader code is Apache-2.0 (repo LICENSE)
- access
- public(pip install seacrowd
- dataset_card
- present(README documents the library, and every dataloader carries a datasheet with license, description and citation)
The loader code is open source, but the corpora it serves come from many publishers, each on its own terms. What a user may do with the data therefore depends on which dataset is loaded, and the library reports that license alongside each one.
- https://huggingface.co/datasets/SEACrowd/nusax_senti/raw/main/README.md recorded 2026-09-24
Loader card for NusaX-Senti: dataset homepage https://github.com/IndoNLP/nusax, "Source: 1.0.0. SEACrowd: 2024.06.20.", license CC BY-SA 4.0
- https://raw.githubusercontent.com/SEACrowd/seacrowd-datahub/master/LICENSE recorded 2026-09-24
Apache License, Version 2.0, January 2004
- https://raw.githubusercontent.com/SEACrowd/seacrowd-datahub/master/README.md recorded 2026-09-24
"`seacrowd` also supports loading the metadata (e.g., license, description, citation, etc.) of the dataloaders"; install with `pip install seacrowd`
Adoption
1 medium confidenceAdoption is counted as PyPI downloads of the `seacrowd` library over the last month. Use through Hugging Face loader mirrors, direct clones of the repository, or the underlying datasets fetched from their own hosts is not captured.
- https://pypistats.org/api/packages/seacrowd/recent recorded 2026-09-24
last_month: 182 downloads of the seacrowd package (pypistats recent)
Capability
3 medium confidenceSEACrowd gives the languages of Southeast Asia a single loader across text, speech and vision-language data, documented in a paper and per-loader datasheets. It packages other groups' datasets rather than collecting new ones, and no outside model or leaderboard naming it was found, where SEA-HELM backs the region's public leaderboard.
- https://arxiv.org/abs/2406.10118 recorded 2026-09-24
"providing standardized corpora in nearly 1,000 SEA languages across three modalities. Through our SEACrowd benchmarks, we assess ... 36 indigenous languages across 13 tasks"
- https://seacrowd.org recorded 2026-09-24
"Compile the first catalog and benchmark for 500+ Southeast Asian datasets"; grassroots-led community of researchers from Southeast Asia
Verified 2026-09-24