AI Potluck
Model components / Dataset Processing Tools

bigcode-dataset

BigCode Project

Preprocessing, quality filtering, decontamination, and PII detection/anonymization for large code corpora. The filtering and PII pipeline behind The Stack and StarCoder training data.

BigCode bigcode-dataset; Apache-2.0 confirmed live 2026-07-21, but maintenance-quiet (last push 2024-08). Built the filtering + PII pipeline behind The Stack / StarCoder.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI)
source
public(bigcode-project/bigcode-dataset)
governance
BigCode
core-gated
ungated

Fully OSI-licensed (Apache-2.0), full source public.

Adoption

2 medium confidence
2.0

~497 GitHub stars; dormant since 2024 but the load-bearing filtering/PII pipeline behind The Stack and StarCoder.

Capability

4 medium confidence
4.0

Solid code-corpus curation coverage; specialized to code, not actively developed.

Unchanged since 2026-07-30 (last edited, not re-checked)