AI Potluck
Back to Gap Map Model components / Language-specific datasets

Sangraha

AI4Bharat
open / Overall score: 3.7

Sangraha is a pretraining text corpus of about 251 billion tokens across 22 Indic languages and English. It joins three parts: text from human-verified websites, OCR of Indic PDFs and transcribed videos and podcasts; perplexity-filtered text from existing multilingual corpora; and English Wikimedia machine-translated into 14 languages and romanized. AI4Bharat built it with Setu, its open cleaning and deduplication pipeline.

Openness

5 high confidence
5.0
license
cc-by-4.0(Hugging Face card metadata
access
public(Hugging Face gated: false)
dataset_card
present(card lists sources, components and per-language token counts)

Sangraha downloads without a gate under CC BY 4.0, and its card, paper and GitHub repository document every source and the cleaning code. The MIT license in that repository applies to the pipeline, not to the text.

Adoption

3 high confidence
3.0

Hugging Face downloads of the single Sangraha repository. A download is one fetch of some files, often a single language split, and says nothing about which models were trained on it.

Capability

4 medium confidence
4.0

Sangraha is the largest open pretraining corpus for Indian languages, covering twenty-two of them plus English at about twelve times the size of IndicCorp, and the IndicBERT models and Soket's Pragna model list it as training data, though about two thirds of its tokens are machine-translated Wikimedia text. Its roughly quarter of a trillion tokens still sit well below the trillions in English web corpora such as FineWeb and DCLM.

Verified 2026-09-24