Samvaad-Hi
Sarvam AISamvaad-Hi is an instruction-tuning dataset of about 100,000 multi-turn conversations in English, Hindi and Hinglish written with an Indian context. Each record is a list of user and assistant messages in chat format. Sarvam AI publishes it, and its one-line card does not say how the conversations were produced.
The card is one sentence, so how the conversations were written or generated is not documented.
Openness
4 high confidence- license
- apache-2.0(card metadata)
- access
- public
- dataset_card
- partial(a single sentence
The conversations are under a permissive license with no gate. What is missing is any account of where they came from, which the one-sentence card does not give.
- https://huggingface.co/api/datasets/sarvamai/samvaad-hi-v1 recorded 2026-09-24
gated: false; cardData license: apache-2.0; 101,476 training examples
- https://huggingface.co/datasets/sarvamai/samvaad-hi-v1/raw/main/README.md recorded 2026-09-24
license: apache-2.0; "100k high-quality conversations in English, Hindi, and Hinglish curated exclusively with an Indic context."
Adoption
1 high confidenceHugging Face downloads of the single repository. Use through models fine-tuned on it is not counted.
- https://huggingface.co/api/datasets/sarvamai/samvaad-hi-v1 recorded 2026-09-24
387 downloads in the trailing 30 days for sarvamai/samvaad-hi-v1
Capability
2 medium confidenceSamvaad-Hi supplies about a hundred thousand Hindi and Hinglish chat conversations, but its card says nothing about how they were made and there is no paper. WangchanThaiInstruct, by comparison, documents that annotators and domain experts wrote and checked every example.
- https://huggingface.co/datasets/sarvamai/samvaad-hi-v1 recorded 2026-09-24
"Models trained or fine-tuned on sarvamai/samvaad-hi-v1": mradermacher/Gaja-v2.00-i1-GGUF, mradermacher/Gaja-v1.00-GGUF and others
Verified 2026-09-24