AI Potluck
Back to Gap Map Organization

SIL Global

foundation · United States

Scores

1 product on the map — 1 open-ish.

Bloom Library datasets

Openness

3 high confidence
3.0
license
mixed-per-subset(Each entry carries its own Creative Commons license: CC-BY, CC-BY-SA, and several NC and ND variants.)
access
auto(Automatic Hugging Face gate
dataset_card
present

Each story keeps the Creative Commons license its author chose, and most of those forbid commercial use or derivatives. The gate approves automatically but shares the user's contact details and adds a pledge against discriminatory use.

Adoption

1 high confidence
1.0

Hugging Face downloads over the trailing month summed across the four Bloom Library repositories. Reading the books on bloomlibrary.org is not counted.

Capability

1 medium confidence
1.0

The datasets gather community-written storybooks in hundreds of languages into text, captioning, storytelling and speech sets documented in an EMNLP paper, and their text went into the BLOOM model's training data as a small part of its corpus. With a couple of stories for most languages they are tiny beside English pretraining corpora of trillions of tokens, and a pooled corpus such as the one behind the Glot language model gathers far more text per language.

  • https://arxiv.org/abs/2210.14712 recorded 2026-09-24

    "covers 363 languages across 32 language families" for language modeling, image captioning, visual storytelling, and speech.

  • https://huggingface.co/datasets/sil-ai/bloom-lm recorded 2026-09-24

    "It includes data from 364 languages across 31 language families. There is a mean of 32 stories and median of 2 stories per language." "this data was used in the training of the BLOOM model".