AI Potluck
Back to Gap Map Model components / Language-specific datasets

Atlaset

AtlasIA
restricted / Overall score: 1.0

Atlaset is a Moroccan Darija text corpus for pretraining and training language models, about 155 million tokens in its training split. It combines dozens of existing sources, including social media comments, news summaries, Wikipedia, YouTube subtitles, and translated instruction sets such as ShareGPT and FLAN, each listed with its row count. Abdelaziz Bounhar curated it, and AtlasIA publishes it.

Openness

2 high confidence
2.0
license
not-clearly-stated-on-card(no license in the card, the API tags, or a LICENSE file
access
auto(automatic approval after agreeing to share contact information)
dataset_card
present(card lists every source with its row count)

Access is granted automatically once a user shares contact details, but no license is given for the compilation or its many sources, so reuse terms are unknown. The card does document where each part came from.

Adoption

1 high confidence
1.0

Hugging Face downloads of the gated repository. The figure cannot show use through the Al-Atlas models trained on it.

Capability

1 medium confidence
1.0

Atlaset gathers most openly available Moroccan Darija text, from social media and subtitles to translated instruction sets, into one documented corpus, and AtlasIA trained its Al-Atlas language model and a Moroccan XLM-RoBERTa on it. At under two hundred million tokens it is tiny beside English pretraining corpora of trillions of tokens, and small even beside naab, which merges billions of words of Farsi.

Verified 2026-09-24