SeaExam
Alibaba DAMO AcademySeaExam is a multiple-choice benchmark from DAMO-NLP-SG, the group behind SeaLLMs, for testing language models in Indonesian, Thai, Vietnamese, Chinese and English. It combines real human exam questions adapted from M3Exam, standardized to four shuffled options, with 2,850 MMLU questions machine-translated using Google Translate. Its companion SeaBench adds open-ended multi-turn and instruction-following questions in Indonesian, Thai and Vietnamese, and both feed the SeaExam leaderboard.
Openness
5 high confidence- license
- apache-2.0(card metadata for both SeaExam and SeaBench)
- access
- public(both repos ungated)
- dataset_card
- present(describes the M3Exam adjustments and the MMLU translation
Both sets download without a gate under Apache-2.0, answers included, and the SeaExam card explains how each part was built. The SeaBench card says little beyond its languages and purpose.
- https://huggingface.co/api/datasets/SeaLLMs/SeaExam recorded 2026-09-24
gated: false; license:apache-2.0
- https://huggingface.co/datasets/SeaLLMs/SeaBench/raw/main/README.md recorded 2026-09-24
license: apache-2.0; "SeaBench evaluates models' multi-turn and instruction-following abilities across Indonesian, Thai, and Vietnamese"
- https://huggingface.co/datasets/SeaLLMs/SeaExam/raw/main/README.md recorded 2026-09-24
license: apache-2.0; "All answers have been mapped to a numerical value"; "translated ... using Google Translate"
Adoption
1 high confidenceAdoption is Hugging Face downloads summed over SeaExam and SeaBench, most of them for SeaExam. Evaluations run through the project's own leaderboard code or local copies are not separately counted.
- https://huggingface.co/api/datasets/SeaLLMs/SeaBench recorded 2026-09-24
87 downloads in the trailing 30 days for SeaLLMs/SeaBench
- https://huggingface.co/api/datasets/SeaLLMs/SeaExam recorded 2026-09-24
447 downloads in the trailing 30 days for SeaLLMs/SeaExam
Capability
3 medium confidenceSeaExam backs a public leaderboard for Indonesian, Thai, Vietnamese, Chinese and English, and its paper appeared at NAACL. Part of it is MMLU run through Google Translate and the rest adapts the earlier M3Exam, and it tests exam knowledge only, where SEA-HELM spans many task types in more languages.
- https://aclanthology.org/2025.findings-naacl.341/ recorded 2026-09-24
SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia, Liu et al.
- https://huggingface.co/datasets/SeaLLMs/SeaExam recorded 2026-09-24
Number of rows: 23,813; Space using SeaLLMs/SeaExam: SeaLLMs/LLM_Leaderboard_for_SEA
- https://huggingface.co/datasets/SeaLLMs/SeaExam/raw/main/README.md recorded 2026-09-24
"The original M3Exam dataset is constructed with real human exam questions"; "We randomly selected 50 questions from each subject, totaling 2850 questions"
- https://raw.githubusercontent.com/DAMO-NLP-SG/SeaExam/main/README.md recorded 2026-09-24
"The leaderboard showcases results from two complementary benchmarks: SeaExam and SeaBench"; cites Findings of NAACL 2025 paper
Verified 2026-09-24