AI Potluck
Back to Gap Map Model components / Language-specific datasets

SEA-Instruct

AI Singapore
gated / Overall score: 1.7

SEA-Instruct is an instruction-tuning dataset from AI Singapore for Southeast Asian languages and contexts, pairing prompts filtered from open data and AI Singapore's own synthetic prompts with model-generated responses. Qwen3 models tagged prompts and drafted answers and DeepSeek-V3.1 revised them, and the release keeps only prompts rated excellent, coherent and natural. It spans English, Chinese and ten regional languages, each row tagged for domain, task, complexity and cultural knowledge.

Openness

3 high confidence
3.0
license
odc-by
access
auto(Hugging Face auto-approved gate
dataset_card
present(pipeline models, schema and tag definitions)

The data is under a permissive license and the card spells out how prompts were filtered and responses generated. Downloading requires accepting a contact-sharing gate on Hugging Face, which is approved automatically.

Adoption

1 high confidence
1.0

Adoption is Hugging Face downloads of the single repository. Training runs typically fetch it once, so the count cannot show how many fine-tunes used it.

Capability

2 medium confidence
2.0

SEA-Instruct covers twelve languages with rows tagged for domain, task and cultural knowledge, and AI Singapore's Nemotron-SEA-LION v4.8 models are trained on it. Every response is written and revised by models, and some prompts are synthetic or translated from English, where WangchanThaiInstruct is written entirely by people.

Verified 2026-09-24