ArabicWeb24
LightOnArabicWeb24 is a pretraining corpus of Arabic web text from a custom crawl, processed with the datatrove library; the card gives more than 28 billion tokens for the fully cleaned and deduplicated version. A second version skips the sentence deduplication step, and the processing code and small ablation models trained on both versions are released alongside. LightOn built it with INSAT.
The card headline says more than 39 billion tokens while its body says more than 28 billion for the cleaned version; the description uses the body figure.
Openness
3 high confidence- license
- odc-by(card metadata and API tag)
- access
- auto(automatic approval after agreeing to share contact information)
- dataset_card
- present(card describes both versions, fields and the processing library)
The corpus carries an attribution-only data license, and access is granted automatically once a user shares contact details on the Hub. The detailed processing steps are in blog posts rather than on the card or in a paper.
- https://huggingface.co/api/datasets/lightonai/ArabicWeb24 recorded 2026-09-24
API JSON: "gated": "auto"; tag license:odc-by.
- https://huggingface.co/datasets/lightonai/ArabicWeb24 recorded 2026-09-24
"You need to agree to share your contact information to access this dataset"; "We are releasing two datasets versions".
Adoption
1 high confidenceHugging Face downloads of the single repository, which holds both versions. Each shard fetched from a large corpus counts separately, so the figure is not a count of training runs.
- https://huggingface.co/api/datasets/lightonai/ArabicWeb24 recorded 2026-09-24
386 downloads in the trailing 30 days for lightonai/ArabicWeb24
Capability
3 medium confidenceArabicWeb24 is a documented Arabic web corpus released with its processing code. No model outside LightOn's own ablations is known to be trained on it, unlike 101 Billion Arabic Words, which ModernAraBERT uses, or naab for Farsi, a larger documented corpus merged from many sources.
- https://huggingface.co/datasets/lightonai/ArabicWeb24 recorded 2026-09-24
"more than 28 billion tokens of cleaned and deduplicated Arabic web data from a customized crawl ... processed using the large scale data processing library datatrove"; "Total file size: 471 GB"; models listed: ArabicWeb24-ablation-model-v1, -v5.
Verified 2026-09-24