AI Potluck
Back to Gap Map Model components / Language-specific datasets

Khayyam Challenge

Raia Center
restricted / Overall score: 3.1

The Khayyam Challenge, also called PersianMMLU, tests language models on 20,192 four-choice questions in Persian from 38 tasks drawn from Persian examinations, from lower primary to upper secondary school. The questions are original rather than translated and carry metadata such as human response rates, difficulty ratings, and descriptive answers. The RAIA Center publishes it.

The Hub README is empty; the description follows the paper and the gating terms.

Openness

2 high confidence
2.0
license
cc-by-nd-4.0(card metadata
access
manual(manual approval after a form sharing contact information and accepting the terms)
dataset_card
no(README.md exists but is empty)

Access needs a form and manual approval, and the terms allow only non-commercial academic research and forbid derivative works, including new benchmarks built from the questions. The repository has no card, so its contents are described only in the paper. The no-derivatives term alone would keep it from being open data, since a model trained on the questions is arguably a derivative.

Adoption

1 high confidence
1.0

Hugging Face downloads of the gated repository. Each approved user fetches it directly, so the figure is closer to a count of users than for ungated sets, but copies inside leaderboards are not counted.

Capability

4 medium confidence
4.0

The Khayyam Challenge is a documented Persian knowledge benchmark of original exam questions with human response rates, and the ParsBench leaderboard uses it. It is a step below ArabicMMLU because it has no card and carries restrictive terms, which make it harder to adopt.

Verified 2026-09-24