AI Potluck
Back to Gap Map Model components / Language-specific datasets

SEA-HELM

AI Singapore
gated / Overall score: 4.4(strong)

SEA-HELM (Southeast Asian Holistic Evaluation of Language Models) is AI Singapore's evaluation suite for large language models in Southeast Asian languages. It groups tasks under five pillars: classic NLP tasks, LLM-specific tasks such as instruction following and multi-turn chat, linguistic diagnostics, culture, and safety. Test sets combine native datasets, human translations and new handcrafted sets such as Kalahi. The suite absorbed the earlier BHASA benchmark and feeds the SEA-HELM leaderboard.

The suite's code lives on GitHub and its test sets on Hugging Face; many of those test sets are samples of existing public datasets.

Openness

3 high confidence
3.0
license
mixed-per-subset(code is MIT
access
auto(task repos use Hugging Face's auto-approved gate with contact sharing)
dataset_card
present(each task card lists sources, licenses, languages and split statistics

The evaluation code is MIT, and the task data is available through automatically approved Hugging Face gates. Each task keeps the license of the dataset it samples, so some subsets are non-commercial or restricted to research use, including the Vietnamese hate-speech source.

Adoption

3 high confidence
3.0

Adoption is Hugging Face downloads summed over the sixteen task repositories. The total counts pulls of the task repositories, so repeated evaluation runs weigh as much as new users, and runs from local copies are not counted.

Capability

5 medium confidence
5.0

SEA-HELM tests eleven Southeast Asian languages across classic NLP tasks, chat, instruction following, linguistics, culture and safety, the widest open suite for the region, and it backs a public leaderboard. Much of its content is sampled from existing datasets such as FLORES, XL-Sum and NusaX.

Verified 2026-09-24