AI Potluck
Back to Gap Map Model components / Language-specific datasets

SeaExam

Alibaba DAMO Academy
open / Overall score: 2.4

SeaExam is a multiple-choice benchmark from DAMO-NLP-SG, the group behind SeaLLMs, for testing language models in Indonesian, Thai, Vietnamese, Chinese and English. It combines real human exam questions adapted from M3Exam, standardized to four shuffled options, with 2,850 MMLU questions machine-translated using Google Translate. Its companion SeaBench adds open-ended multi-turn and instruction-following questions in Indonesian, Thai and Vietnamese, and both feed the SeaExam leaderboard.

Openness

5 high confidence
5.0
license
apache-2.0(card metadata for both SeaExam and SeaBench)
access
public(both repos ungated)
dataset_card
present(describes the M3Exam adjustments and the MMLU translation

Both sets download without a gate under Apache-2.0, answers included, and the SeaExam card explains how each part was built. The SeaBench card says little beyond its languages and purpose.

Adoption

1 high confidence
1.0

Adoption is Hugging Face downloads summed over SeaExam and SeaBench, most of them for SeaExam. Evaluations run through the project's own leaderboard code or local copies are not separately counted.

Capability

3 medium confidence
3.0

SeaExam backs a public leaderboard for Indonesian, Thai, Vietnamese, Chinese and English, and its paper appeared at NAACL. Part of it is MMLU run through Google Translate and the rest adapts the earlier M3Exam, and it tests exam knowledge only, where SEA-HELM spans many task types in more languages.

Verified 2026-09-24