101 Billion Arabic Words
ClusterlabAiThe 101 Billion Arabic Words Dataset is a pretraining corpus of Arabic web text, mixing Modern Standard Arabic and dialects, for training and fine-tuning language models. Arabic content was extracted from Common Crawl WET files, then cleaned and deduplicated, yielding about 101 billion words in 33 million documents. It is curated by the Clusterlab team.
Openness
5 high confidence- license
- apache-2.0(card metadata and the card body ("License: Apache 2.0"))
- access
- public(ungated on the Hugging Face Hub)
- dataset_card
- present(card and paper describe the Common Crawl extraction and cleaning)
The full corpus downloads from the Hub without a sign-up, and its card and paper describe how the web text was extracted and cleaned. The text itself is crawled web content, so the terms of the original pages still sit behind the permissive grant on the compilation.
- https://huggingface.co/api/datasets/ClusterlabAi/101_billion_arabic_words_dataset recorded 2026-09-24
API JSON: "gated": false, "private": false; tag license:apache-2.0.
- https://huggingface.co/datasets/ClusterlabAi/101_billion_arabic_words_dataset/raw/main/README.md recorded 2026-09-24
Card metadata "license: apache-2.0"; body: "License: Apache 2.0"; "gathered data from specified sources, primarily Common Crawl, and extracted Arabic content from WET files".
Adoption
2 high confidenceHugging Face downloads of the single Hub repository. Downloads of a corpus this size mostly count partial fetches of its shards and cannot show how many training runs used it.
- https://huggingface.co/api/datasets/ClusterlabAi/101_billion_arabic_words_dataset recorded 2026-09-24
1904 downloads in the trailing 30 days for ClusterlabAi/101_billion_arabic_words_dataset
Capability
4 medium confidenceThe corpus gives Arabic a large cleaned and deduplicated Common Crawl collection, documented in a paper whose authors call it the largest Arabic dataset available, and ModernAraBERT is trained on it. At about a hundred billion words it is among the largest single-language corpora outside English and Chinese, though still short of the trillions of tokens in FineWeb or DCLM, and like Sangraha it is a big documented corpus with little outside use recorded.
- https://arxiv.org/abs/2405.01590 recorded 2026-09-24
Abstract: "extracting a substantial volume of text from the Common Crawl WET files ... The result is the 101 Billion Arabic Words Dataset, the largest Arabic dataset available to date".
- https://huggingface.co/datasets/ClusterlabAi/101_billion_arabic_words_dataset recorded 2026-09-24
Hub page: "Number of rows: 33,059,988"; "Models trained or fine-tuned on" lists gizadatateam/ModernAraBERT.
Verified 2026-09-24