Sarvam-2T
Sarvam AISarvam-2T is Sarvam AI's Indic pretraining corpus of about 2 trillion tokens across ten Indian languages: Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil and Telugu. Most of it was produced with synthetic data generation rather than web crawling, and the languages are split almost evenly except for Hindi. Sarvam used it to train its Sarvam-1 model.
Openness
1 high confidence- license
- not-clearly-stated-on-card(Not distributed
- access
- closed(Described in the Sarvam-1 post
- dataset_card
- no(No card
Sarvam released the Sarvam-1 model trained on this corpus but not the corpus itself.
- https://www.sarvam.ai/blogs/sarvam-1 recorded 2026-09-24
"Our training corpus, which we call Sarvam-2T, encompasses ~2 trillion Indic tokens in total"; "The model can be downloaded from 🤗 Hub"; no corpus download offered.
Adoption
not assessedThe corpus was never distributed, so there is no download or usage figure to read.
- https://www.sarvam.ai/blogs/sarvam-1 recorded 2026-09-24
The Sarvam-1 post describes Sarvam-2T as the training corpus and links no download of it, so no usage figure exists
Capability
2 medium confidenceBy Sarvam's account the corpus is very large and trained its Sarvam-1 model, but the text is largely synthetic and is described only in a blog post, with no paper, card or data release beyond a few snippets. IndicCorp v2 is smaller but is documented in a paper and trained IndicBERT v2.
- https://www.sarvam.ai/blogs/sarvam-1 recorded 2026-09-24
"we have developed a high-quality training corpus of 2 trillion tokens, specifically for 10 Indic languages"; "Hindi, which comprises about 20% of the data"; "through advanced synthetic-data-generation techniques".
Verified 2026-09-24