AI Potluck
Model components / Dataset Processing Tools

bigcode-dataset

BigCode Project

Preprocessing, quality filtering, decontamination, and PII detection/anonymization for large code corpora. The filtering and PII pipeline behind The Stack and StarCoder training data.

BigCode bigcode-dataset, maintenance-quiet since August 2024. Built the filtering and PII pipeline behind The Stack and StarCoder. Verified 2026-08-13 via the GitHub README and the LICENSE body.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI)
source
public(bigcode-project/bigcode-dataset)
governance
BigCode
core-gated
ungated

Fully OSI-licensed (Apache-2.0), full source public.

Adoption

1 medium confidence
1.0

497 GitHub stars on bigcode-project/bigcode-dataset, which bands at <1K stars, level 1 on the stars fallback scale. That scale caps at level 3, because a star is not a use. The repo has been dormant since 2024. As the pipeline behind The Stack and StarCoder it matters more than its star count suggests, but importance is not adoption, and no download, install or customer figure is published for this product, so stars remain the only honest signal and the level is directional.

Capability

4 medium confidence
4.0

Solid code-corpus curation coverage; specialized to code, not actively developed.

Verified 2026-08-12