IndicGenBench
Google ResearchIndicGenBench is a benchmark of generation tasks for large language models in 29 Indic languages across 13 scripts, covering cross-lingual summarization, machine translation, and multilingual and cross-lingual question answering. It extends CrossSum, FLORES, XQuAD and XorQA by having people translate their English examples into each target language, which makes the test data multi-way parallel. Google Research releases it as four task sets.
Openness
2 high confidence- license
- cc-by-sa-4.0(Flores-IN and XQuAD-IN)+mit(XorQA-IN)+cc-by-nc-sa-4.0(CrossSum-IN)
- access
- public(all four Hugging Face repositories gated: false
- dataset_card
- present(GitHub README and per-task cards describe tasks, fields and splits)
Each task set keeps the license of the dataset it extends, so the summarization set is non-commercial while the translation and question-answering sets allow commercial use. All four download without a gate.
- https://huggingface.co/api/datasets/google/IndicGenBench_crosssum_in recorded 2026-09-24
"gated": false; license tag cc-by-nc-sa-4.0.
- https://huggingface.co/api/datasets/google/IndicGenBench_flores_in recorded 2026-09-24
"gated": false; license tag cc-by-sa-4.0.
- https://huggingface.co/api/datasets/google/IndicGenBench_xorqa_in recorded 2026-09-24
"gated": false; license tag mit.
- https://huggingface.co/api/datasets/google/IndicGenBench_xquad_in recorded 2026-09-24
"gated": false; license tag cc-by-sa-4.0.
- https://raw.githubusercontent.com/google-research-datasets/indic-gen-bench/main/LICENSE recorded 2026-09-24
"The different datasets distributed as part of this repository are subject to different licenses based on the license of the source dataset they were derived from."
- https://raw.githubusercontent.com/google-research-datasets/indic-gen-bench/main/README.md recorded 2026-09-24
"XQuAD-IN, Flores-IN: Under CC BY-SA 4.0 license ... XorQA-IN: Under MIT license ... CrossSum-IN: Under CC BY-NC-SA 4.0 license".
Adoption
2 high confidenceHugging Face downloads summed over the four task repositories. The same data is also cloned from GitHub, which publishes no count.
- https://huggingface.co/api/datasets/google/IndicGenBench_crosssum_in recorded 2026-09-24
711 downloads in the trailing 30 days for google/IndicGenBench_crosssum_in
- https://huggingface.co/api/datasets/google/IndicGenBench_flores_in recorded 2026-09-24
1093 downloads in the trailing 30 days for google/IndicGenBench_flores_in
- https://huggingface.co/api/datasets/google/IndicGenBench_xorqa_in recorded 2026-09-24
465 downloads in the trailing 30 days for google/IndicGenBench_xorqa_in
- https://huggingface.co/api/datasets/google/IndicGenBench_xquad_in recorded 2026-09-24
802 downloads in the trailing 30 days for google/IndicGenBench_xquad_in
Capability
3 medium confidenceIndicGenBench is documented and reaches 29 Indic languages with human translations. No leaderboard or model report built on it outside its own paper was found, whereas IndicXTREME was released with the IndicBERT v2 models it evaluates.
- https://arxiv.org/abs/2404.16816 recorded 2026-09-24
Abstract: "the largest benchmark for evaluating LLMs on user-facing generation tasks across a diverse set 29 of Indic languages covering 13 scripts".
Verified 2026-09-24