Speech and Language Processing Lab, Sharif University of Technology
lab · IranScores
1 product on the map — 1 open.
Openness
5 medium confidence- license
- mit(card metadata and API tag
- access
- public(both repositories ungated)
- dataset_card
- present(card lists each source corpus and the cleaning process)
The corpus, its raw version and the cleaning code are all downloadable, and the card names every source it merged. The license line on the card ends in a question mark, and the merged sources carry their own terms, so the grant is less certain than the metadata tag suggests.
- https://huggingface.co/api/datasets/SLPL/naab recorded 2026-09-24
API JSON: "gated": false, "private": false; tag license:mit.
- https://huggingface.co/datasets/SLPL/naab/raw/main/README.md recorded 2026-09-24
Card metadata license "mit"; "### Licensing Information\n\nmit?"; source sections for Persian NLP, AGP, OSCAR-fa, Telegram and LSCP.
Adoption
2 high confidenceHugging Face downloads summed over the cleaned and raw repositories, most of them for the cleaned one. Shard fetches of a large corpus are counted individually, so the figure is not a count of training runs.
- https://huggingface.co/api/datasets/SLPL/naab recorded 2026-09-24
2560 downloads in the trailing 30 days for SLPL/naab
- https://huggingface.co/api/datasets/SLPL/naab-raw recorded 2026-09-24
155 downloads in the trailing 30 days for SLPL/naab-raw
Capability
3 medium confidencenaab is the largest cleaned public Farsi pretraining corpus, compiled from web and messaging sources rather than newly collected and documented in a paper with its raw version and preprocessing code, and a Farsi text-to-text transformer and a Persian ALBERT are trained on it. At about fifteen billion words it is a sizable national corpus, yet small beside the trillions of tokens in English pretraining sets such as FineWeb.
- https://arxiv.org/abs/2208.13486 recorded 2026-09-24
Abstract: "the largest publicly available, cleaned, and ready-to-use Farsi textual corpus. naab consists of 130GB of data, comprising over 250 million paragraphs and 15 billion words".
- https://huggingface.co/datasets/SLPL/naab recorded 2026-09-24
Hub page "Models trained or fine-tuned on SLPL/naab" lists shekar-ai/albert-base-v2-persian-zwnj-naab-mlm and SLPL/t5-fa.