AtlasIA
unknownScores
1 product on the map — 1 closed.
Openness
2 high confidence- license
- not-clearly-stated-on-card(no license in the card, the API tags, or a LICENSE file
- access
- auto(automatic approval after agreeing to share contact information)
- dataset_card
- present(card lists every source with its row count)
Access is granted automatically once a user shares contact details, but no license is given for the compilation or its many sources, so reuse terms are unknown. The card does document where each part came from.
- https://huggingface.co/api/datasets/atlasia/Atlaset recorded 2026-09-24
API JSON: "gated": "auto"; no license tag in cardData or tags.
- https://huggingface.co/datasets/atlasia/Atlaset recorded 2026-09-24
"You need to agree to share your contact information to access this dataset"; card lists training and test sources with counts; no license shown.
Adoption
1 high confidenceHugging Face downloads of the gated repository. The figure cannot show use through the Al-Atlas models trained on it.
- https://huggingface.co/api/datasets/atlasia/Atlaset recorded 2026-09-24
60 downloads in the trailing 30 days for atlasia/Atlaset
Capability
1 medium confidenceAtlaset gathers most openly available Moroccan Darija text, from social media and subtitles to translated instruction sets, into one documented corpus, and AtlasIA trained its Al-Atlas language model and a Moroccan XLM-RoBERTa on it. At under two hundred million tokens it is tiny beside English pretraining corpora of trillions of tokens, and small even beside naab, which merges billions of words of Farsi.
- https://huggingface.co/datasets/atlasia/Atlaset recorded 2026-09-24
"For the training set: 155,501,098 tokens"; "It combines various sources"; models listed: atlasia/Al-Atlas-0.5B, atlasia/XLM-RoBERTa-Morocco.