naab
Speech and Language Processing Lab, Sharif University of Technologynaab is a cleaned Farsi text corpus for pretraining language models, with about 130 GB, 250 million paragraphs, and 15 billion words. It merges existing Persian corpora, OSCAR-fa, the formerly private AGP corpus, Telegram channels, and LSCP, and a raw version (naab-raw) and the preprocessing code are published beside it. Sharif University of Technology's Speech and Language Processing Lab builds it.
Openness
5 medium confidence- license
- mit(card metadata and API tag
- access
- public(both repositories ungated)
- dataset_card
- present(card lists each source corpus and the cleaning process)
The corpus, its raw version and the cleaning code are all downloadable, and the card names every source it merged. The license line on the card ends in a question mark, and the merged sources carry their own terms, so the grant is less certain than the metadata tag suggests.
- https://huggingface.co/api/datasets/SLPL/naab recorded 2026-09-24
API JSON: "gated": false, "private": false; tag license:mit.
- https://huggingface.co/datasets/SLPL/naab/raw/main/README.md recorded 2026-09-24
Card metadata license "mit"; "### Licensing Information\n\nmit?"; source sections for Persian NLP, AGP, OSCAR-fa, Telegram and LSCP.
Adoption
2 high confidenceHugging Face downloads summed over the cleaned and raw repositories, most of them for the cleaned one. Shard fetches of a large corpus are counted individually, so the figure is not a count of training runs.
- https://huggingface.co/api/datasets/SLPL/naab recorded 2026-09-24
2560 downloads in the trailing 30 days for SLPL/naab
- https://huggingface.co/api/datasets/SLPL/naab-raw recorded 2026-09-24
155 downloads in the trailing 30 days for SLPL/naab-raw
Capability
3 medium confidencenaab is the largest cleaned public Farsi pretraining corpus, compiled from web and messaging sources rather than newly collected and documented in a paper with its raw version and preprocessing code, and a Farsi text-to-text transformer and a Persian ALBERT are trained on it. At about fifteen billion words it is a sizable national corpus, yet small beside the trillions of tokens in English pretraining sets such as FineWeb.
- https://arxiv.org/abs/2208.13486 recorded 2026-09-24
Abstract: "the largest publicly available, cleaned, and ready-to-use Farsi textual corpus. naab consists of 130GB of data, comprising over 250 million paragraphs and 15 billion words".
- https://huggingface.co/datasets/SLPL/naab recorded 2026-09-24
Hub page "Models trained or fine-tuned on SLPL/naab" lists shekar-ai/albert-base-v2-persian-zwnj-naab-mlm and SLPL/t5-fa.
Verified 2026-09-24