CIDAR
ARBMLCIDAR (Culturally Relevant Instruction Dataset for Arabic) holds 10,000 Arabic instruction and response pairs for tuning language models. About 9,109 come from the Alpagasus set translated into Arabic with ChatGPT and about 891 are native Arabic grammar questions from Al Jazeera's Ask the Teacher site; about 12 reviewers corrected every pair for cultural fit. ARBML publishes it with two 100-item evaluation sets.
The two companion evaluation sets, CIDAR-EVAL-100 and CIDAR-MCQ-100, are part of the same release and are counted with the main set.
Openness
5 high confidence- license
- apache-2.0(card metadata and body
- access
- public(all three Hub repositories ungated)
- dataset_card
- present(card describes sources, translation and review)
The data, the annotation app and the paper are all public under a permissive license. The card asks that it be used for research only, but that request is not a term of the license.
- https://huggingface.co/api/datasets/arbml/CIDAR recorded 2026-09-24
API JSON: "gated": false, "private": false.
- https://huggingface.co/datasets/arbml/CIDAR/raw/main/README.md recorded 2026-09-24
Card metadata "license: apache-2.0"; body: "CIDAR is intended for research purposes only."
- https://raw.githubusercontent.com/ARBML/CIDAR/main/LICENSE recorded 2026-09-24
Apache License, Version 2.0, January 2004, full text.
- https://raw.githubusercontent.com/ARBML/CIDAR/main/README.md recorded 2026-09-24
"CIDAR is released under Apache 2.0."
Adoption
1 high confidenceHugging Face downloads summed over the instruction set and its two evaluation sets, nearly all of them for the instruction set. Downloads do not show which of them fed a released model.
- https://huggingface.co/api/datasets/arbml/CIDAR recorded 2026-09-24
386 downloads in the trailing 30 days for arbml/CIDAR
- https://huggingface.co/api/datasets/arbml/cidar-eval-100 recorded 2026-09-24
35 downloads in the trailing 30 days for arbml/CIDAR-EVAL-100
- https://huggingface.co/api/datasets/arbml/cidar-mcq-100 recorded 2026-09-24
33 downloads in the trailing 30 days for arbml/CIDAR-MCQ-100
Capability
2 medium confidenceCIDAR is documented and every pair was reviewed by people for cultural fit, but most of it began as machine translation from English. Only small community fine-tunes use it, whereas WangchanThaiInstruct was written in Thai by people from the start.
- https://arxiv.org/abs/2402.03177 recorded 2026-09-24
Abstract: "the first open Arabic instruction-tuning dataset culturally-aligned by human reviewers. CIDAR contains 10,000 instruction and output pairs".
- https://huggingface.co/datasets/arbml/CIDAR recorded 2026-09-24
Hub page "Models trained or fine-tuned on arbml/CIDAR" lists VohoAI/voho-saudi-chat-4b and Ruqiya/Fine-Tuning-Gemma-2b-it-for-Arabic.
- https://huggingface.co/datasets/arbml/CIDAR/raw/main/README.md recorded 2026-09-24
"selecting around 9,109 samples from Alpagasus dataset then translating it to Arabic using ChatGPT ... around 891 Arabic grammar instructions ... reviewed by around 12 reviewers."
Verified 2026-09-24