UberText
lang-ukUberText 2.0 is a corpus of modern Ukrainian of about 3.27 billion tokens in 8.59 million texts, spanning news from 38 websites, fiction, public Telegram channels, the Ukrainian Wikipedia and Supreme Court decisions. It is offered for direct download in several processing layers, from cleaned text to sentence-split and tokenized files. Dmytro Chaplynskyi and the lang-uk community build it.
Openness
2 medium confidence- license
- not-clearly-stated-on-card(No license on the UberText 2.0 page, the lang.org.ua corpora card, or the UNLP 2023 paper
- access
- public(Direct download links on lang.org.ua, no registration.)
- dataset_card
- present(Corpus page with composition and download layers
The corpus downloads directly with no registration, but neither its page nor its paper names a license; the paper argues that the use of the texts falls within fair use and Ukrainian copyright law. The maintainers offer to remove texts at a rights holder's request.
- https://lang.org.ua/en/corpora/ recorded 2026-09-24
Corpora card: "A 3.27-billion-token corpus of modern Ukrainian: news, Wikipedia, fiction, court decisions and social media"; no license field for UberText 2.0.
- https://lang.org.ua/en/ubertext/ recorded 2026-09-24
Composition of five subcorpora and download links for base, cleansed, sentence-split and tokenized layers; no license statement.
- https://lang.org.ua/static/downloads/preprints/ubertext_paper.pdf recorded 2026-09-24
"We believe that our use of these texts falls within the bounds of fair use and Ukrainian copyright law ... we are willing to remove any texts from our corpus upon request"
Adoption
not assessedNo download figure could be read. The corpus is published as files on lang.org.ua, which shows no download count, and there is no Hugging Face repository or package to count instead.
- https://lang.org.ua/en/ubertext/ recorded 2026-09-24
The corpus page lists direct download links for each subcorpus and shows no download counter
Capability
2 medium confidenceUberText is the large corpus of modern Ukrainian, documented at the Ukrainian NLP workshop, and lang-uk lists Flair and fastText embeddings and Electra and GPT models built on it, while the Malyuk pretraining corpus folds it in, though its named uses are smaller models rather than a large language model. At about three billion tokens it is a small fraction of the trillions in English pretraining corpora.
- https://aclanthology.org/2023.unlp-1.1/ recorded 2026-09-24
"It has 3.274 billion tokens, consists of 8.59 million texts and takes up 32 gigabytes of space."
- https://lang.org.ua/en/corpora/ recorded 2026-09-24
"Malyuk — Ukrainian LM pre-training corpus. Combined corpus of Ukrainian texts (UberText 2.0, Oscar, Wikipedia and more)"
- https://lang.org.ua/en/ubertext/ recorded 2026-09-24
"Projects using UberText 2.0": flair embeddings, fastText vectors, GPT-2 models of different sizes, Electra embeddings trained on UberText 2.0.
Verified 2026-09-24