IndicAlign
AI4BharatIndicAlign is an instruction-tuning and safety-alignment collection of about 74.7 million prompt-response pairs in 14 Indic languages. It pools translated and romanized English sets such as Dolly and OpenAssistant, instruction pairs built from IndoWordNet entries, crowd-sourced Anudesh prompts, Wikipedia-grounded conversations generated with Llama 2, and toxic prompts paired with refusals. AI4Bharat released it with Sangraha as part of IndicLLMSuite.
Openness
5 medium confidence- license
- cc-by-4.0(Hugging Face card metadata)
- access
- public(Hugging Face gated: false)
- dataset_card
- present(card describes each of the ten components with sources and statistics)
IndicAlign downloads without a gate under CC BY 4.0, and the card names the source of each component. Many responses were generated by Llama 2 70B Chat and Mistral 7B, and the card does not say whether those models' terms carry over to the data.
- https://huggingface.co/api/datasets/ai4bharat/indic-align recorded 2026-09-24
"gated": false; cardData license "cc-by-4.0".
- https://huggingface.co/datasets/ai4bharat/indic-align recorded 2026-09-24
Card header "License: cc-by-4.0"; components list IndicShareLlama, Dolly-T, OpenAssistant-T, WikiHow, IndoWordNet, Anudesh, Wiki-Conv, Wiki-Chat, HHRLHF-T, Toxic-Matrix, with Llama2-70B-Chat and Mistral-7B named as generators.
Adoption
2 high confidenceHugging Face downloads of the IndicAlign repository, which holds all ten components. A download does not show whether the data went into a released model.
- https://huggingface.co/api/datasets/ai4bharat/indic-align recorded 2026-09-24
2063 downloads in the trailing 30 days for ai4bharat/indic-align
Capability
2 medium confidenceIndicAlign is well documented in the IndicLLMSuite paper, but nearly all of its pairs are short IndoWordNet dictionary entries recast as instructions, and most of the rest are translated or model-generated. Only small community fine-tunes list it as training data, whereas WangchanThaiInstruct is written and checked by people.
- https://huggingface.co/datasets/ai4bharat/indic-align recorded 2026-09-24
Statistics table: IndoWordNet 74,272.2k examples, Wiki-Chat 202k, Wiki-Conv 144k, Toxic Matrix 90.3k, other components 15k to 37k each.
- https://raw.githubusercontent.com/AI4Bharat/IndicLLMSuite/master/README.md recorded 2026-09-24
"IndicAlign is the largest multilingual Instruction Fine-tuning dataset for Indic Languages, comprising ... around 74.7 million prompt-response pairs."
Verified 2026-09-24