Updesh
MicrosoftUpdesh is a synthetic instruction-tuning dataset of about 9.5 million examples in 13 Indian languages and English. Its reasoning part translates Orca-AgentInstruct and OrcaMath tasks with Llama 3.1 405B, and its generative part builds multi-hop questions, dialogues, summaries and similar tasks from language-specific Wikipedia with Qwen3-235B. Microsoft Research built it to post-train smaller models for Indian languages.
Openness
2 high confidence- license
- Microsoft-Research-License(LICENSE.md in the repository: non-commercial research use only
- access
- public(Hugging Face gated: false)
- dataset_card
- present(card describes sources, generator models, per-subset volumes and quality checks)
Updesh downloads without a gate, but its Microsoft Research license limits use to non-commercial research and forbids distributing the data or modified versions of it. The card adds that the data is not intended for commercial deployment. Either restriction alone keeps it from being open data.
- https://huggingface.co/api/datasets/microsoft/Updesh_beta recorded 2026-09-24
"gated": false; cardData license "other".
- https://huggingface.co/datasets/microsoft/Updesh_beta recorded 2026-09-24
"We release this data under the Microsoft Research License." Limitations: "Updesh is released for research purposes only and is not intended for commercial deployment".
- https://huggingface.co/datasets/microsoft/Updesh_beta/raw/main/LICENSE.md recorded 2026-09-24
"use the Materials solely for non-commercial, non-revenue generating, research purposes ... Data: ... you may not distribute the data or your modifications to the data."
Adoption
2 high confidenceHugging Face downloads of the single Updesh repository. Because the license bars redistribution, copies should not circulate elsewhere, but a download still does not show use in a model.
- https://huggingface.co/api/datasets/microsoft/Updesh_beta recorded 2026-09-24
8186 downloads in the trailing 30 days for microsoft/Updesh_beta
Capability
2 medium confidenceUpdesh is carefully documented and checked with human quality ratings, but every example is machine-translated or generated by large models, and the only models trained on it are the ones in its own paper. WangchanThaiInstruct, by contrast, is written by annotators and domain experts.
- https://arxiv.org/abs/2509.21294 recorded 2026-09-24
Abstract: "a high-quality large-scale synthetic instruction-following dataset comprising 9.5M data points across 13 Indian languages and English ... 10K human assessments".
- https://huggingface.co/datasets/microsoft/Updesh_beta recorded 2026-09-24
"Reasoning Data: ~6.8M translated tuples; Generative Data: ~2.1M synthesized tuples"; translation model Llama-3.1-405B-Instruct; generation model Qwen3-235B-A22B.
Verified 2026-09-24