AI Potluck
Back to Gap Map Model components / Language-specific datasets

CIDAR

ARBML
open / Overall score: 1.7

CIDAR (Culturally Relevant Instruction Dataset for Arabic) holds 10,000 Arabic instruction and response pairs for tuning language models. About 9,109 come from the Alpagasus set translated into Arabic with ChatGPT and about 891 are native Arabic grammar questions from Al Jazeera's Ask the Teacher site; about 12 reviewers corrected every pair for cultural fit. ARBML publishes it with two 100-item evaluation sets.

The two companion evaluation sets, CIDAR-EVAL-100 and CIDAR-MCQ-100, are part of the same release and are counted with the main set.

Openness

5 high confidence
5.0
license
apache-2.0(card metadata and body
access
public(all three Hub repositories ungated)
dataset_card
present(card describes sources, translation and review)

The data, the annotation app and the paper are all public under a permissive license. The card asks that it be used for research only, but that request is not a term of the license.

Adoption

1 high confidence
1.0

Hugging Face downloads summed over the instruction set and its two evaluation sets, nearly all of them for the instruction set. Downloads do not show which of them fed a released model.

Capability

2 medium confidence
2.0

CIDAR is documented and every pair was reviewed by people for cultural fit, but most of it began as machine translation from English. Only small community fine-tunes use it, whereas WangchanThaiInstruct was written in Thai by people from the start.

Verified 2026-09-24