Masader
ARBMLMasader is a public catalog of Arabic text and speech datasets, describing each with more than two dozen attributes such as dialect, domain, source, annotation method and license, and linking to where its publisher hosts it. New entries are submitted through a form and added by the maintainers. It began in the BigScience workshop and is maintained by the ARBML community.
Openness
5 high confidence- license
- GPL-3.0(OSI)
- source
- public(github.com/ARBML/masader)
- core features withheld
- no — volunteer research community project
The catalog site and its data are published under GPL-3.0 and build from the repository with Jekyll, and the web service it calls is MIT-licensed. It is a volunteer research project with nothing sold.
- https://raw.githubusercontent.com/ARBML/masader-webservice/HEAD/LICENSE recorded 2026-10-07
"MIT License Copyright (c) 2022 ARBML" for the catalogue's web service.
- https://raw.githubusercontent.com/ARBML/masader/HEAD/CONTRIBUTING.md recorded 2026-10-07
"We use jekyll"; "Run the site locally with bundle exec jekyll serve".
- https://raw.githubusercontent.com/ARBML/masader/HEAD/license recorded 2026-10-07
GNU General Public License, Version 3, 29 June 2007, full text.
- https://raw.githubusercontent.com/ARBML/masader/HEAD/README.md recorded 2026-10-07
"Masader was developed in 2021 as part of the BigScience project for open research"; "further developed by the arbml team and community."
Adoption
1 low confidenceMasader is a website built from its repository and publishes no visitor or download count, so GitHub stars are the only signal. A star records interest in the catalog rather than a use of it.
- https://img.shields.io/github/stars/ARBML/masader.json recorded 2026-10-07
{"label":"stars","message":"208"} for ARBML/masader, read 2026-10-07.
Capability
1 medium confidenceMasader describes Arabic text and speech datasets in detail, from dialect and domain to license, and links each one to where its publisher hosts it. It stores none of the data, while the UCI repository hosts the datasets donated to it and serves them by ID.
- https://arxiv.org/abs/2110.06744 recorded 2026-10-07
"the largest public catalogue for Arabic NLP datasets, which consists of 200 datasets annotated with 25 attributes" (2021 paper).
- https://raw.githubusercontent.com/ARBML/masader/HEAD/README.md recorded 2026-10-07
"The first online catalogue for Arabic NLP datasets"; "Link: direct link to the dataset or instructions on how to download it"; "If you want to add a new dataset, use this form."
Verified 2026-10-07, except some scores, still awaiting re-confirmation