AI Potluck
Back to Gap Map Model components / Language-specific datasets

IndicGenBench

Google Research
restricted / Overall score: 2.7

IndicGenBench is a benchmark of generation tasks for large language models in 29 Indic languages across 13 scripts, covering cross-lingual summarization, machine translation, and multilingual and cross-lingual question answering. It extends CrossSum, FLORES, XQuAD and XorQA by having people translate their English examples into each target language, which makes the test data multi-way parallel. Google Research releases it as four task sets.

Openness

2 high confidence
2.0
license
cc-by-sa-4.0(Flores-IN and XQuAD-IN)+mit(XorQA-IN)+cc-by-nc-sa-4.0(CrossSum-IN)
access
public(all four Hugging Face repositories gated: false
dataset_card
present(GitHub README and per-task cards describe tasks, fields and splits)

Each task set keeps the license of the dataset it extends, so the summarization set is non-commercial while the translation and question-answering sets allow commercial use. All four download without a gate.

Adoption

2 high confidence
2.0

Hugging Face downloads summed over the four task repositories. The same data is also cloned from GitHub, which publishes no count.

Capability

3 medium confidence
3.0

IndicGenBench is documented and reaches 29 Indic languages with human translations. No leaderboard or model report built on it outside its own paper was found, whereas IndicXTREME was released with the IndicBERT v2 models it evaluates.

  • https://arxiv.org/abs/2404.16816 recorded 2026-09-24

    Abstract: "the largest benchmark for evaluating LLMs on user-facing generation tasks across a diverse set 29 of Indic languages covering 13 scripts".

Verified 2026-09-24