AI Potluck
Back to Gap Map Model components / Language-specific datasets

Cendol Collection

IndoNLP
open / Overall score: 3.1

Cendol Collection is the instruction-tuning data IndoNLP used to train the Cendol family of language models for Indonesian and the country's local languages. Version 1 holds task-specific prompts built from NLP datasets for sentiment analysis, translation, summarization and question answering. Version 2 adds general-knowledge and human-centric prompts used for the Cendol-Chat models. Both are described in the Cendol paper.

Both dataset cards are copies of the Cendol model card and say little about how the prompts were assembled; the description relies on the model card's account of the two training stages.

Openness

4 medium confidence
4.0
license
apache-2.0(card metadata on both repos
access
public(both repos ungated)
dataset_card
partial(cards reproduce the model card and name the subset but do not describe sources or composition)

Both versions download without a gate under Apache-2.0. Their cards describe the Cendol models rather than the prompts, so how the data was assembled has to be read from the paper.

Adoption

1 high confidence
1.0

Adoption is Hugging Face downloads summed over the v1 and v2 repositories. Fine-tunes that reuse a local copy are not counted.

Capability

4 medium confidence
4.0

The collection trained IndoNLP's Cendol Instruct and Chat models, from 300 million to 13 billion parameters, and the Cendol paper documents it. Like WangchanThaiInstruct it has a paper and named models trained on it, but its task prompts are rebuilt from existing labeled datasets and its cards do not describe the contents.

Verified 2026-09-24