AI Potluck
Back to Gap Map Model components / Language-specific datasets

WangchanX-FLAN

VISTEC
open / Overall score: 2.4

WangchanX-FLAN is a FLAN-style instruction-tuning mixture centered on Thai, combining about 3.6 million examples converted from more than two dozen existing datasets for summarization, translation, question answering, classification and text generation. Sources range from ThaiSum, the scb-mt-en-th-2020 parallel corpus and Wisesight sentiment to English sets such as UltraChat, FLAN v2 and xP3x. VISTEC-depa AI Research Institute of Thailand assembled it, and the build script is public.

Openness

4 high confidence
4.0
license
mixed-per-subset(card metadata says 'other' with a LICENSE link, but the LICENSE file is empty
access
public(ungated)
dataset_card
present(per-source table of origin, size, task, domain and license)

The mixture downloads without a gate, and each source dataset's license is recorded in the card table and on every row. The release adds no license of its own, since its LICENSE file is empty.

Adoption

1 high confidence
1.0

Adoption is Hugging Face downloads of the v6.1 repository. Earlier versions of the mixture and local rebuilds from the public script are not counted.

Capability

3 medium confidence
3.0

WangchanX-FLAN lists every source and license, and the WangchanLION v2 instruct model is trained on it. All of it is reformatted from existing datasets, some of them machine-generated, which places it below WangchanThaiInstruct, whose Thai examples annotators and domain experts wrote from scratch.

Verified 2026-09-24