AI Potluck
Back to Gap Map Model components / Language-specific datasets

SMOL

Google
open / Overall score: 1.3

SMOL is a Google suite of professionally translated training data for machine translation into more than 200 low-resource languages. It combines SmolSent, sentences picked for broad vocabulary coverage; SmolDoc, whole documents chosen for topic coverage, with factuality ratings; and the GATITOS word and phrase lexicon. Volunteers have added translations for further languages since the first release.

Openness

5 high confidence
5.0
license
cc-by-4.0
access
public
dataset_card
present

The translations carry an attribution-only license and download from Hugging Face without a gate.

Adoption

2 high confidence
2.0

Hugging Face downloads over the trailing month for the one SMOL repository, which also carries GATITOS.

Capability

1 medium confidence
1.0

SMOL gives more than two hundred low-resource languages professionally translated text, documented in two papers, though no model or benchmark beyond its own experiments is known to be built on it. At a few million translated tokens it is minute beside the billions of pairs in the largest parallel collections, and smaller even than the IIT Bombay English-Hindi corpus for one language pair.

Verified 2026-09-24