SEA-Instruct
AI SingaporeSEA-Instruct is an instruction-tuning dataset from AI Singapore for Southeast Asian languages and contexts, pairing prompts filtered from open data and AI Singapore's own synthetic prompts with model-generated responses. Qwen3 models tagged prompts and drafted answers and DeepSeek-V3.1 revised them, and the release keeps only prompts rated excellent, coherent and natural. It spans English, Chinese and ten regional languages, each row tagged for domain, task, complexity and cultural knowledge.
Openness
3 high confidence- license
- odc-by
- access
- auto(Hugging Face auto-approved gate
- dataset_card
- present(pipeline models, schema and tag definitions)
The data is under a permissive license and the card spells out how prompts were filtered and responses generated. Downloading requires accepting a contact-sharing gate on Hugging Face, which is approved automatically.
- https://huggingface.co/api/datasets/aisingapore/SEA-Instruct-2602 recorded 2026-09-24
gated: auto; license:odc-by
- https://huggingface.co/datasets/aisingapore/SEA-Instruct-2602 recorded 2026-09-24
"You need to agree to share your contact information to access this dataset"; "License: ODC-By"; pipeline: Qwen3-235B-A22B tagging, Qwen3-32B responses, DeepSeek-V3.1 revision
Adoption
1 high confidenceAdoption is Hugging Face downloads of the single repository. Training runs typically fetch it once, so the count cannot show how many fine-tunes used it.
- https://huggingface.co/api/datasets/aisingapore/SEA-Instruct-2602 recorded 2026-09-24
289 downloads in the trailing 30 days for aisingapore/SEA-Instruct-2602
Capability
2 medium confidenceSEA-Instruct covers twelve languages with rows tagged for domain, task and cultural knowledge, and AI Singapore's Nemotron-SEA-LION v4.8 models are trained on it. Every response is written and revised by models, and some prompts are synthetic or translated from English, where WangchanThaiInstruct is written entirely by people.
- https://huggingface.co/api/datasets/aisingapore/SEA-Instruct-2602 recorded 2026-09-24
size_categories:1M<n<10M; languages en, id, vi, th, ta, tl, ms, my, km, lo, zh
- https://huggingface.co/datasets/aisingapore/SEA-Instruct-2602 recorded 2026-09-24
"Models trained or fine-tuned on aisingapore/SEA-Instruct-2602: aisingapore/Nemotron-SEA-LION-v4.8-120B-A12B ..."; Total file size: 16.3 GB
Verified 2026-09-24