AI Potluck
Back to Gap Map Model components / Language-specific datasets

WangchanThaiInstruct

VISTEC
open / Overall score: 3.1

WangchanThaiInstruct is a human-authored Thai instruction dataset for both instruction tuning and evaluation, covering four professional domains (medical, finance, retail and legal) and seven task types from summarization to multiple-choice QA. Annotators, domain experts and researchers wrote and checked it through a multi-stage quality-control process, and it has about 40,000 rows split into train and test sets. It is published by the VISTEC-depa AI Research Institute of Thailand.

Openness

4 high confidence
4.0
license
mixed-per-subset(card metadata says cc-by-sa-4.0, but the card states rows carry NC or SA licenses according to their source, recorded in a per-row License column)
access
public(ungated)
dataset_card
present(domains, tasks, subsets and paper link)

The data downloads without a gate, but its license varies row by row: some rows allow commercial use under share-alike terms and others are non-commercial. Anyone reusing it has to filter on the per-row license field.

Adoption

1 high confidence
1.0

Adoption is Hugging Face downloads of the single repository. It serves as both a training set and a test set, and the count does not distinguish the two uses.

Capability

4 medium confidence
4.0

WangchanThaiInstruct is written entirely by Thai annotators and domain experts rather than translated or generated, it is documented in an EMNLP paper, and models such as WangchanLION v2 are trained on it. It covers four professional domains in one language, and the models the Hub lists as trained on it include quantized copies of WangchanLION v2 and a Llama 3.1 Thai fine-tune.

Verified 2026-09-24