Ungoliant
OSCAR ProjectRust pipeline that generates the OSCAR corpus from Common Crawl WET files: language identification, filtering, and metadata annotation. The official OSCAR / Colossal-OSCAR construction pipeline.
OSCAR project Ungoliant, still receiving commits (last push November 2025). The pipeline that produces the OSCAR / Colossal-OSCAR multilingual corpora. Verified 2026-08-13 via the GitHub README and the LICENSE body.
Openness
5 high confidence- license
- Apache-2.0(OSI)
- source
- public(oscar-project/ungoliant)
- governance
- OSCAR project
- core-gated
- ungated
Fully OSI-licensed (Apache-2.0), full source public.
- https://github.com/oscar-project/ungoliant/blob/main/LICENSE recorded 2026-08-13
Apache License Version 2.0 text
- https://raw.githubusercontent.com/oscar-project/ungoliant/main/README.md recorded 2026-08-13
README documents the whole pipeline building and installing from the published repo; no paid tier, enterprise edition or license key appears anywhere in it, so nothing is withheld from the source
Adoption
1 medium confidence178 GitHub stars on oscar-project/ungoliant, which bands at <1K stars, level 1 on the stars fallback scale. That scale caps at level 3, because a star is not a use. As the official pipeline behind the OSCAR corpora it matters more than its star count suggests, but importance is not adoption, and no download, install or customer figure is published for this product, so stars remain the only honest signal and the level is directional.
- https://api.github.com/repos/oscar-project/ungoliant recorded 2026-08-12
stargazers_count = 178 for oscar-project/ungoliant
Capability
3 medium confidenceFocused single-corpus pipeline; mature for its scope.
- https://raw.githubusercontent.com/oscar-project/ungoliant/main/README.md recorded 2026-08-13
README still describes the OSCAR generation pipeline over Common Crawl, with language identification, filtering and an optional KenLM feature
Verified 2026-08-12