SMOL
GoogleSMOL is a Google suite of professionally translated training data for machine translation into more than 200 low-resource languages. It combines SmolSent, sentences picked for broad vocabulary coverage; SmolDoc, whole documents chosen for topic coverage, with factuality ratings; and the GATITOS word and phrase lexicon. Volunteers have added translations for further languages since the first release.
Openness
5 high confidence- license
- cc-by-4.0
- access
- public
- dataset_card
- present
The translations carry an attribution-only license and download from Hugging Face without a gate.
- https://huggingface.co/api/datasets/google/smol recorded 2026-09-24
gated: false; license tag cc-by-4.0.
- https://huggingface.co/datasets/google/smol/raw/main/README.md recorded 2026-09-24
Card front matter: "license: cc-by-4.0"; describes SmolDoc, SmolSent, GATITOS and the factuality annotations.
Adoption
2 high confidenceHugging Face downloads over the trailing month for the one SMOL repository, which also carries GATITOS.
- https://huggingface.co/api/datasets/google/smol recorded 2026-09-24
1829 downloads in the trailing 30 days for google/smol
Capability
1 medium confidenceSMOL gives more than two hundred low-resource languages professionally translated text, documented in two papers, though no model or benchmark beyond its own experiments is known to be built on it. At a few million translated tokens it is minute beside the billions of pairs in the largest parallel collections, and smaller even than the IIT Bombay English-Hindi corpus for one language pair.
- https://arxiv.org/abs/2502.12301 recorded 2026-09-24
"SMOL has been translated into 124 (and growing) under-resourced languages ... for a total of 6.1M translated tokens."
- https://huggingface.co/datasets/google/smol/raw/main/README.md recorded 2026-09-24
"a collection professional translations into 221 Low-Resource Languages"; "the first 113 SMOL languages were commissioned by professional translation vendors".
Verified 2026-09-24