AI Potluck
Back to Gap Map Model components / Language-specific datasets

TurkishMMLU

Arda Yüksel
restricted / Overall score: 2.4

TurkishMMLU tests language models on multiple-choice questions from Turkish high-school curricula, covering nine subjects in natural sciences, mathematics, Turkish language and literature, and social sciences and humanities. Curriculum experts wrote the more than 10,000 questions, and each carries a correctness ratio as a difficulty signal. Arda Yüksel, Abdullatif Köksal and colleagues publish it with evaluation code.

The Hub repository holds dev and test splits of 1,845 rows; the full question set is available by emailing the authors, as the README states.

Openness

2 medium confidence
2.0
license
not-clearly-stated-on-card(no license in the Hub card metadata or API tags
access
public(the Hub subset is ungated
dataset_card
present(card gives the abstract, subjects and evaluation results)

Only a subset of the questions is openly downloadable, and the rest must be requested from the authors by email. No license is stated anywhere, so reuse terms for either part are unknown.

Adoption

1 high confidence
1.0

Hugging Face downloads of the public subset. Copies of the full set obtained by email are not counted, and evaluation harnesses fetch the benchmark on each run.

Capability

3 medium confidence
3.0

TurkishMMLU is a documented benchmark of expert-written Turkish exam questions, and a community Turkish LLM leaderboard runs on it. Only a small part is openly posted, which limits its use as a shared test, whereas ArabicMMLU publishes its full question set.

Verified 2026-09-24