AI Potluck
Back to Gap Map Organization

lang-uk

unknown

Scores

1 product on the map — 1 closed.

UberText

Openness

2 medium confidence
2.0
license
not-clearly-stated-on-card(No license on the UberText 2.0 page, the lang.org.ua corpora card, or the UNLP 2023 paper
access
public(Direct download links on lang.org.ua, no registration.)
dataset_card
present(Corpus page with composition and download layers

The corpus downloads directly with no registration, but neither its page nor its paper names a license; the paper argues that the use of the texts falls within fair use and Ukrainian copyright law. The maintainers offer to remove texts at a rights holder's request.

  • https://lang.org.ua/en/corpora/ recorded 2026-09-24

    Corpora card: "A 3.27-billion-token corpus of modern Ukrainian: news, Wikipedia, fiction, court decisions and social media"; no license field for UberText 2.0.

  • https://lang.org.ua/en/ubertext/ recorded 2026-09-24

    Composition of five subcorpora and download links for base, cleansed, sentence-split and tokenized layers; no license statement.

  • https://lang.org.ua/static/downloads/preprints/ubertext_paper.pdf recorded 2026-09-24

    "We believe that our use of these texts falls within the bounds of fair use and Ukrainian copyright law ... we are willing to remove any texts from our corpus upon request"

Adoption

not assessed

No download figure could be read. The corpus is published as files on lang.org.ua, which shows no download count, and there is no Hugging Face repository or package to count instead.

Capability

2 medium confidence
2.0

UberText is the large corpus of modern Ukrainian, documented at the Ukrainian NLP workshop, and lang-uk lists Flair and fastText embeddings and Electra and GPT models built on it, while the Malyuk pretraining corpus folds it in, though its named uses are smaller models rather than a large language model. At about three billion tokens it is a small fraction of the trillions in English pretraining corpora.

  • https://aclanthology.org/2023.unlp-1.1/ recorded 2026-09-24

    "It has 3.274 billion tokens, consists of 8.59 million texts and takes up 32 gigabytes of space."

  • https://lang.org.ua/en/corpora/ recorded 2026-09-24

    "Malyuk — Ukrainian LM pre-training corpus. Combined corpus of Ukrainian texts (UberText 2.0, Oscar, Wikipedia and more)"

  • https://lang.org.ua/en/ubertext/ recorded 2026-09-24

    "Projects using UberText 2.0": flair embeddings, fastText vectors, GPT-2 models of different sizes, Electra embeddings trained on UberText 2.0.