AI Potluck
Back to Gap Map Model components / Language-specific datasets

FLEURS

Google
open / Overall score: 4.7(leading)

FLEURS is an n-way parallel speech benchmark in 102 languages, made by recording people reading the FLORES sentences aloud. Each language has roughly ten hours of training speech plus separate development and test speakers, and it supports speech recognition, spoken language identification, speech translation and retrieval. Google built it as the speech counterpart of FLORES.

Openness

5 high confidence
5.0
license
cc-by-4.0
access
public
dataset_card
present

The recordings carry an attribution-only license and download from Hugging Face without a gate.

Adoption

4 high confidence
4.0

Hugging Face downloads over the trailing month for the FLEURS repository. Copies pulled through evaluation toolkits that mirror it are not counted.

Capability

5 medium confidence
5.0

FLEURS gives 102 languages the same read sentences, so speech recognition, language identification and speech translation can be compared across all of them. It is documented in its paper, and Meta's MMS work reports its word error rates against Whisper on it.

  • https://arxiv.org/abs/2205.12446 recorded 2026-09-24

    "an n-way parallel speech dataset in 102 languages ... with approximately 12 hours of speech supervision per language".

  • https://arxiv.org/abs/2305.13516 recorded 2026-09-24

    "more than halves the word error rate of Whisper on 54 languages of the FLEURS benchmark".

  • https://huggingface.co/datasets/google/fleurs recorded 2026-09-24

    "We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision."

Verified 2026-09-24