Croissant
MLCommonsCroissant is MLCommons' metadata format for machine-learning datasets, combining a dataset's metadata, file descriptions, record structure and default ML semantics in one JSON-LD file built on schema.org. The mlcroissant Python library validates and loads Croissant files, and a Responsible AI extension records provenance, collection and bias information. Major dataset hubs publish Croissant descriptions.
Scored as the format and its reference library, not the hubs that serve it. Tagged self-attested: a validator checks the structure, not the truth of what it records.
Openness
5 high confidence- license
- Apache-2.0(OSI
- source
- public
- core features withheld
- no
The specification and library are published under Apache-2.0 by MLCommons and developed in the open, with no enterprise, commercial or proprietary directory in the repository. One contributed subdirectory is MIT, which changes nothing.
- https://cdn.jsdelivr.net/gh/mlcommons/croissant@main/croissant-rdf/LICENSE recorded 2026-09-26
"MIT License Copyright (c) 2024 David Steinberg" for the croissant-rdf subdirectory
- https://cdn.jsdelivr.net/gh/mlcommons/croissant@main/LICENSE.md recorded 2026-09-26
Apache License, Version 2.0, full text
- https://cdn.jsdelivr.net/gh/mlcommons/croissant@main/README.md recorded 2026-09-26
"Croissant is currently under development by the community. You can try the Croissant implementation, mlcroissant"
- https://ungh.cc/repos/mlcommons/croissant/files/main recorded 2026-09-26
Full file listing of mlcommons/croissant main, 782 paths; no ee/, enterprise/, commercial/ or proprietary/ directory
Adoption
2 medium confidencePyPI downloads of mlcroissant, the reference library, measure use of the format's tooling. Hubs that emit Croissant files without the library are not counted.
- https://pypistats.org/api/packages/mlcroissant/recent recorded 2026-09-26
last_month downloads = 28,529 for mlcroissant
Capability
3 high confidenceA Croissant file follows a public schema that anyone can validate with open tooling, so a reader can check its structure without the publisher. Its provenance and responsible-AI fields are still the publisher's own account, two steps below NVIDIA's hardware attestation.
- https://cdn.jsdelivr.net/gh/mlcommons/croissant@main/docs/croissant-rai-spec.md recorded 2026-09-26
Croissant Responsible AI (RAI) vocabulary specification
- https://cdn.jsdelivr.net/gh/mlcommons/croissant@main/health/croissant-validator-neurips/README.md recorded 2026-09-26
Croissant validator for NeurIPS dataset submissions
- https://cdn.jsdelivr.net/gh/mlcommons/croissant@main/README.md recorded 2026-09-26
"Croissant is a high-level format for machine learning datasets that combines metadata, resource file descriptions, data structure, and default ML semantics into a single file"; builds on schema.org
Verified 2026-09-26