AI Potluck
Back to Gap Map Model components / Language-specific datasets

IndicAlign

AI4Bharat
open / Overall score: 2.0

IndicAlign is an instruction-tuning and safety-alignment collection of about 74.7 million prompt-response pairs in 14 Indic languages. It pools translated and romanized English sets such as Dolly and OpenAssistant, instruction pairs built from IndoWordNet entries, crowd-sourced Anudesh prompts, Wikipedia-grounded conversations generated with Llama 2, and toxic prompts paired with refusals. AI4Bharat released it with Sangraha as part of IndicLLMSuite.

Openness

5 medium confidence
5.0
license
cc-by-4.0(Hugging Face card metadata)
access
public(Hugging Face gated: false)
dataset_card
present(card describes each of the ten components with sources and statistics)

IndicAlign downloads without a gate under CC BY 4.0, and the card names the source of each component. Many responses were generated by Llama 2 70B Chat and Mistral 7B, and the card does not say whether those models' terms carry over to the data.

Adoption

2 high confidence
2.0

Hugging Face downloads of the IndicAlign repository, which holds all ten components. A download does not show whether the data went into a released model.

Capability

2 medium confidence
2.0

IndicAlign is well documented in the IndicLLMSuite paper, but nearly all of its pairs are short IndoWordNet dictionary entries recast as instructions, and most of the rest are translated or model-generated. Only small community fine-tunes list it as training data, whereas WangchanThaiInstruct is written and checked by people.

Verified 2026-09-24