AI Potluck
Back to Gap Map Model components / Language-specific datasets

Korpus Malti

Maltese Language Resource Server, University of Malta
restricted / Overall score: 1.0

Korpus Malti is a general corpus of written Maltese spanning 19 genres, from parliamentary debates, court proceedings and laws to the press, blogs, theses, literature and Wikipedia. Version 4.0, the data used to train the BERTu model, holds about 466 million tokens in 20.8 million sentences. The Maltese Language Resource Server team at the University of Malta curates it.

Openness

2 high confidence
2.0
license
cc-by-nc-sa-4.0
access
auto(Hugging Face gated: auto
dataset_card
present(Describes each of the 19 genre subsets and versioning

The files sit behind an automatic Hugging Face gate that asks for contact details, and the license rules out commercial use. The literary subset cannot be distributed in full, since its texts are included by permission of their copyright holders.

Adoption

1 high confidence
1.0

Hugging Face downloads of the dataset repository over the trailing 30 days. A download is a file fetch rather than a trained model, and copies mirrored elsewhere are not counted.

Capability

1 medium confidence
1.0

Korpus Malti is the general corpus of written Maltese, drawn from nineteen genres including parliament, courts, the press and literature, and the paper that introduced it trained the BERTu and mBERTu models on it. At under half a billion tokens it is smaller than national collections such as Danish Dynaword and very far from the trillions of tokens in English pretraining corpora.

Verified 2026-09-24