SEA-HELM
AI SingaporeSEA-HELM (Southeast Asian Holistic Evaluation of Language Models) is AI Singapore's evaluation suite for large language models in Southeast Asian languages. It groups tasks under five pillars: classic NLP tasks, LLM-specific tasks such as instruction following and multi-turn chat, linguistic diagnostics, culture, and safety. Test sets combine native datasets, human translations and new handcrafted sets such as Kalahi. The suite absorbed the earlier BHASA benchmark and feeds the SEA-HELM leaderboard.
The suite's code lives on GitHub and its test sets on Hugging Face; many of those test sets are samples of existing public datasets.
Openness
3 high confidence- license
- mixed-per-subset(code is MIT
- access
- auto(task repos use Hugging Face's auto-approved gate with contact sharing)
- dataset_card
- present(each task card lists sources, licenses, languages and split statistics
The evaluation code is MIT, and the task data is available through automatically approved Hugging Face gates. Each task keeps the license of the dataset it samples, so some subsets are non-commercial or restricted to research use, including the Vietnamese hate-speech source.
- https://huggingface.co/collections/aisingapore/sea-helm-evaluation-datasets recorded 2026-09-24
Collection 'SEA-HELM Evaluation Datasets' listing the 14 task repositories
- https://huggingface.co/datasets/aisingapore/NLG-Machine-Translation recorded 2026-09-24
"You need to agree to share your contact information to access this dataset"; "It is part of the SEA-HELM leaderboard from AI Singapore"
- https://raw.githubusercontent.com/aisingapore/SEA-HELM/main/docs/datasets_and_prompts.md recorded 2026-09-24
Per-task tables of dataset, nativeness, domain and license, e.g. ViHSD "Research purposes only", XL-Sum "CC BY-NC-SA 4.0"
- https://raw.githubusercontent.com/aisingapore/SEA-HELM/main/LICENSE recorded 2026-09-24
MIT License, Copyright (c) 2025 AI Singapore
- https://raw.githubusercontent.com/aisingapore/SEA-HELM/main/README.md recorded 2026-09-24
"The codes for SEA-HELM is licensed under the MIT license. All datasets are licensed under their respective licenses."
Adoption
3 high confidenceAdoption is Hugging Face downloads summed over the sixteen task repositories. The total counts pulls of the task repositories, so repeated evaluation runs weigh as much as new users, and runs from local copies are not counted.
- https://huggingface.co/api/datasets/aisingapore/Cultural-Evaluation-Kalahi recorded 2026-09-24
947 downloads in the trailing 30 days for aisingapore/Cultural-Evaluation-Kalahi
- https://huggingface.co/api/datasets/aisingapore/Cultural-Evaluation-Kalahi-Judge recorded 2026-09-24
888 downloads in the trailing 30 days for aisingapore/Cultural-Evaluation-Kalahi-Judge
- https://huggingface.co/api/datasets/aisingapore/Instruction-Following-IFEval recorded 2026-09-24
2904 downloads in the trailing 30 days for aisingapore/Instruction-Following-IFEval
- https://huggingface.co/api/datasets/aisingapore/Linguistic-Diagnostics-Pragmatics recorded 2026-09-24
1438 downloads in the trailing 30 days for aisingapore/Linguistic-Diagnostics-Pragmatics
- https://huggingface.co/api/datasets/aisingapore/Linguistic-Diagnostics-Syntax recorded 2026-09-24
1398 downloads in the trailing 30 days for aisingapore/Linguistic-Diagnostics-Syntax
- https://huggingface.co/api/datasets/aisingapore/MultiTurn-Chat-MT-Bench recorded 2026-09-24
80 downloads in the trailing 30 days for aisingapore/MultiTurn-Chat-MT-Bench
- https://huggingface.co/api/datasets/aisingapore/MultiTurn-Chat-MT-Bench-Judge recorded 2026-09-24
2270 downloads in the trailing 30 days for aisingapore/MultiTurn-Chat-MT-Bench-Judge
- https://huggingface.co/api/datasets/aisingapore/NLG-Abstractive-Summarization recorded 2026-09-24
2313 downloads in the trailing 30 days for aisingapore/NLG-Abstractive-Summarization
- https://huggingface.co/api/datasets/aisingapore/NLG-Machine-Translation recorded 2026-09-24
2893 downloads in the trailing 30 days for aisingapore/NLG-Machine-Translation
- https://huggingface.co/api/datasets/aisingapore/NLR-Causal-Reasoning recorded 2026-09-24
2428 downloads in the trailing 30 days for aisingapore/NLR-Causal-Reasoning
- https://huggingface.co/api/datasets/aisingapore/NLR-NLI recorded 2026-09-24
2399 downloads in the trailing 30 days for aisingapore/NLR-NLI
- https://huggingface.co/api/datasets/aisingapore/NLU-Belebele-MCQA recorded 2026-09-24
1588 downloads in the trailing 30 days for aisingapore/NLU-Belebele-MCQA
- https://huggingface.co/api/datasets/aisingapore/NLU-Metaphor recorded 2026-09-24
921 downloads in the trailing 30 days for aisingapore/NLU-Metaphor
- https://huggingface.co/api/datasets/aisingapore/NLU-Question-Answering recorded 2026-09-24
1860 downloads in the trailing 30 days for aisingapore/NLU-Question-Answering
- https://huggingface.co/api/datasets/aisingapore/NLU-Sentiment-Analysis recorded 2026-09-24
2599 downloads in the trailing 30 days for aisingapore/NLU-Sentiment-Analysis
- https://huggingface.co/api/datasets/aisingapore/Safety-Toxicity-Detection recorded 2026-09-24
2568 downloads in the trailing 30 days for aisingapore/Safety-Toxicity-Detection
Capability
5 medium confidenceSEA-HELM tests eleven Southeast Asian languages across classic NLP tasks, chat, instruction following, linguistics, culture and safety, the widest open suite for the region, and it backs a public leaderboard. Much of its content is sampled from existing datasets such as FLORES, XL-Sum and NusaX.
- https://arxiv.org/abs/2502.14301 recorded 2026-09-24
arXiv 2502.14301, SEA-HELM: Southeast Asian Holistic Evaluation of Language Models
- https://raw.githubusercontent.com/aisingapore/SEA-HELM/main/docs/datasets_and_prompts.md recorded 2026-09-24
Task tables name FLORES, XL-Sum, NusaX, XNLI, Belebele, Kalahi, LINDSEA, SEA-IFEval, SEA MT-Bench, Thai Exam
- https://raw.githubusercontent.com/aisingapore/SEA-HELM/main/README.md recorded 2026-09-24
"This suite consist of 5 core pillars: NLP Classics, LLM-specifics, SEA Linguistics, SEA Culture, Safety"; "BHASA has been integrated into SEA-HELM"
Verified 2026-09-24