AI Potluck
Back to Gap Map Model components / Language-specific datasets

Samvaad-Hi

Sarvam AI
open / Overall score: 1.7

Samvaad-Hi is an instruction-tuning dataset of about 100,000 multi-turn conversations in English, Hindi and Hinglish written with an Indian context. Each record is a list of user and assistant messages in chat format. Sarvam AI publishes it, and its one-line card does not say how the conversations were produced.

The card is one sentence, so how the conversations were written or generated is not documented.

Openness

4 high confidence
4.0
license
apache-2.0(card metadata)
access
public
dataset_card
partial(a single sentence

The conversations are under a permissive license with no gate. What is missing is any account of where they came from, which the one-sentence card does not give.

Adoption

1 high confidence
1.0

Hugging Face downloads of the single repository. Use through models fine-tuned on it is not counted.

Capability

2 medium confidence
2.0

Samvaad-Hi supplies about a hundred thousand Hindi and Hinglish chat conversations, but its card says nothing about how they were made and there is no paper. WangchanThaiInstruct, by comparison, documents that annotators and domain experts wrote and checked every example.

Verified 2026-09-24