AI Potluck
Back to Gap Map Model components / Language-specific datasets

SEACrowd

SEACrowd
open / Overall score: 2.4

SEACrowd collects and standardizes datasets for Southeast Asian languages behind one Python library, `seacrowd`, whose dataloaders return text, speech and vision-language corpora in shared schemas with license and citation metadata attached. The hub covers nearly 1,000 languages of the region across three modalities and ships benchmark bundles such as SEACrowd-VL. It is a grassroots research community effort modeled on NusaCrowd.

Scored as the hub and its library. The datasets it loads keep their own publishers and licenses, and the Hugging Face mirrors under the SEACrowd account are loaders for those datasets rather than new releases.

Openness

4 high confidence
4.0
license
mixed-per-subset(The loader code is Apache-2.0 (repo LICENSE)
access
public(pip install seacrowd
dataset_card
present(README documents the library, and every dataloader carries a datasheet with license, description and citation)

The loader code is open source, but the corpora it serves come from many publishers, each on its own terms. What a user may do with the data therefore depends on which dataset is loaded, and the library reports that license alongside each one.

Adoption

1 medium confidence
1.0

Adoption is counted as PyPI downloads of the `seacrowd` library over the last month. Use through Hugging Face loader mirrors, direct clones of the repository, or the underlying datasets fetched from their own hosts is not captured.

Capability

3 medium confidence
3.0

SEACrowd gives the languages of Southeast Asia a single loader across text, speech and vision-language data, documented in a paper and per-loader datasheets. It packages other groups' datasets rather than collecting new ones, and no outside model or leaderboard naming it was found, where SEA-HELM backs the region's public leaderboard.

  • https://arxiv.org/abs/2406.10118 recorded 2026-09-24

    "providing standardized corpora in nearly 1,000 SEA languages across three modalities. Through our SEACrowd benchmarks, we assess ... 36 indigenous languages across 13 tasks"

  • https://seacrowd.org recorded 2026-09-24

    "Compile the first catalog and benchmark for 500+ Southeast Asian datasets"; grassroots-led community of researchers from Southeast Asia

Verified 2026-09-24