Data-Juicer
Alibaba CloudData processing system for foundation models: 200+ composable operators for cleaning, deduplicating, synthesizing, and analyzing text and multimodal LLM training data.
Alibaba Data-Juicer (datajuicer/data-juicer, formerly modelscope); PyPI package py-data-juicer. Verified 2026-08-13 via the GitHub README and the LICENSE body.
Openness
5 high confidence- license
- Apache-2.0(OSI)
- source
- public
- core-gated
- ungated
Fully OSI-licensed (Apache-2.0), full source public.
- https://github.com/datajuicer/data-juicer/blob/main/LICENSE recorded 2026-08-13
Apache License Version 2.0 text
- https://raw.githubusercontent.com/datajuicer/data-juicer/main/README.md recorded 2026-08-13
README documents the whole pipeline building and installing from the published repo; no paid tier, enterprise edition or license key appears anywhere in it, so nothing is withheld from the source
Adoption
1 high confidence2,962 downloads in the trailing 30 days for the declared `py-data-juicer` package, which bands at <10K, level 1 on the software and model adoption scale.
- https://pypistats.org/api/packages/py-data-juicer/recent recorded 2026-08-13
2,962 downloads in the trailing 30 days for py-data-juicer
Capability
4 medium confidenceBroadest operator coverage among the open tools; multimodal. Benchmark participation but no single frontier-topping corpus of its own.
- https://raw.githubusercontent.com/datajuicer/data-juicer/main/README.md recorded 2026-08-13
README still carries the 200+ operators badge and covers cleaning, deduplication, synthesis and analysis across text and multimodal data
Verified 2026-08-13