Korpus Malti
Maltese Language Resource Server, University of MaltaKorpus Malti is a general corpus of written Maltese spanning 19 genres, from parliamentary debates, court proceedings and laws to the press, blogs, theses, literature and Wikipedia. Version 4.0, the data used to train the BERTu model, holds about 466 million tokens in 20.8 million sentences. The Maltese Language Resource Server team at the University of Malta curates it.
Openness
2 high confidence- license
- cc-by-nc-sa-4.0
- access
- auto(Hugging Face gated: auto
- dataset_card
- present(Describes each of the 19 genre subsets and versioning
The files sit behind an automatic Hugging Face gate that asks for contact details, and the license rules out commercial use. The literary subset cannot be distributed in full, since its texts are included by permission of their copyright holders.
- https://huggingface.co/api/datasets/MLRS/korpus_malti recorded 2026-09-24
"gated":"auto", license cc-by-nc-sa-4.0
- https://huggingface.co/datasets/MLRS/korpus_malti recorded 2026-09-24
"You need to agree to share your contact information to access this dataset"; "This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License."
Adoption
1 high confidenceHugging Face downloads of the dataset repository over the trailing 30 days. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted.
- https://huggingface.co/api/datasets/MLRS/korpus_malti recorded 2026-09-24
34 downloads in the trailing 30 days for MLRS/korpus_malti
Capability
1 medium confidenceKorpus Malti is the general corpus of written Maltese, drawn from nineteen genres including parliament, courts, the press and literature, and the paper that introduced it trained the BERTu and mBERTu models on it. At under half a billion tokens it is smaller than national collections such as Danish Dynaword and very far from the trillions of tokens in English pretraining corpora.
- https://arxiv.org/abs/2205.10517 recorded 2026-09-24
"We pre-train and compare two models on the new corpus: a monolingual BERT model trained from scratch (BERTu), and a further pre-trained multilingual BERT (mBERTu)."
- https://arxiv.org/pdf/2205.10517 recorded 2026-09-24
Table 1, Korpus Malti v4.0 corpus distribution: all subsets 131,429 documents, 20,758,071 sentences, 466,601,373 tokens, 2.52 GB.
- https://huggingface.co/api/models?filter=dataset:MLRS/korpus_malti recorded 2026-09-24
Models tagged with the dataset: MLRS/BERTu, MLRS/mBERTu, Cabbache/Fredu-1.7B-Instruct.
Verified 2026-09-24