Sangraha
AI4BharatSangraha is a pretraining text corpus of about 251 billion tokens across 22 Indic languages and English. It joins three parts: text from human-verified websites, OCR of Indic PDFs and transcribed videos and podcasts; perplexity-filtered text from existing multilingual corpora; and English Wikimedia machine-translated into 14 languages and romanized. AI4Bharat built it with Setu, its open cleaning and deduplication pipeline.
Openness
5 high confidence- license
- cc-by-4.0(Hugging Face card metadata
- access
- public(Hugging Face gated: false)
- dataset_card
- present(card lists sources, components and per-language token counts)
Sangraha downloads without a gate under CC BY 4.0, and its card, paper and GitHub repository document every source and the cleaning code. The MIT license in that repository applies to the pipeline, not to the text.
- https://huggingface.co/api/datasets/ai4bharat/sangraha recorded 2026-09-24
"gated": false; cardData license "cc-by-4.0".
- https://huggingface.co/datasets/ai4bharat/sangraha recorded 2026-09-24
Card header "License: cc-by-4.0"; describes Verified, Unverified and Synthetic components and a per-language token table totaling 251,321.0 million.
- https://raw.githubusercontent.com/AI4Bharat/IndicLLMSuite/master/LICENSE recorded 2026-09-24
MIT License, Copyright (c) AI4Bharat, for the repository code.
- https://raw.githubusercontent.com/AI4Bharat/IndicLLMSuite/master/README.md recorded 2026-09-24
README: "We open-source our pre-training dataset Sangraha" and "all the code and other resources used for curating these datasets, including Setu".
Adoption
3 high confidenceHugging Face downloads of the single Sangraha repository. A download is one fetch of some files, often a single language split, and says nothing about which models were trained on it.
- https://huggingface.co/api/datasets/ai4bharat/sangraha recorded 2026-09-24
40034 downloads in the trailing 30 days for ai4bharat/sangraha
Capability
4 medium confidenceSangraha is the largest open pretraining corpus for Indian languages, covering twenty-two of them plus English at about twelve times the size of IndicCorp, and the IndicBERT models and Soket's Pragna model list it as training data, though about two thirds of its tokens are machine-translated Wikimedia text. Its roughly quarter of a trillion tokens still sit well below the trillions in English web corpora such as FineWeb and DCLM.
- https://arxiv.org/abs/2212.05409 recorded 2026-09-24
IndicCorp abstract: "20.9B tokens covering 24 languages".
- https://arxiv.org/abs/2403.06350 recorded 2026-09-24
Abstract: "covering 22 languages, containing a total of 251B tokens and 74.8M instruction-response pairs".
- https://huggingface.co/datasets/ai4bharat/sangraha recorded 2026-09-24
Token table: Verified 64,306.1M, Synthetic 162,707.9M, Unverified 24,307.7M, total 251,321.0M; "Models trained or fine-tuned on" lists IndicBERT-v3-270M, Cadence, pragna-1b.
Verified 2026-09-24