AI Potluck
Back to Gap Map Model components / Language-specific datasets

Updesh

Microsoft
restricted / Overall score: 2.0

Updesh is a synthetic instruction-tuning dataset of about 9.5 million examples in 13 Indian languages and English. Its reasoning part translates Orca-AgentInstruct and OrcaMath tasks with Llama 3.1 405B, and its generative part builds multi-hop questions, dialogues, summaries and similar tasks from language-specific Wikipedia with Qwen3-235B. Microsoft Research built it to post-train smaller models for Indian languages.

Openness

2 high confidence
2.0
license
Microsoft-Research-License(LICENSE.md in the repository: non-commercial research use only
access
public(Hugging Face gated: false)
dataset_card
present(card describes sources, generator models, per-subset volumes and quality checks)

Updesh downloads without a gate, but its Microsoft Research license limits use to non-commercial research and forbids distributing the data or modified versions of it. The card adds that the data is not intended for commercial deployment. Either restriction alone keeps it from being open data.

Adoption

2 high confidence
2.0

Hugging Face downloads of the single Updesh repository. Because the license bars redistribution, copies should not circulate elsewhere, but a download still does not show use in a model.

Capability

2 medium confidence
2.0

Updesh is carefully documented and checked with human quality ratings, but every example is machine-translated or generated by large models, and the only models trained on it are the ones in its own paper. WangchanThaiInstruct, by contrast, is written by annotators and domain experts.

  • https://arxiv.org/abs/2509.21294 recorded 2026-09-24

    Abstract: "a high-quality large-scale synthetic instruction-following dataset comprising 9.5M data points across 13 Indian languages and English ... 10K human assessments".

  • https://huggingface.co/datasets/microsoft/Updesh_beta recorded 2026-09-24

    "Reasoning Data: ~6.8M translated tuples; Generative Data: ~2.1M synthesized tuples"; translation model Llama-3.1-405B-Instruct; generation model Qwen3-235B-A22B.

Verified 2026-09-24