AI Potluck
Back to Gap Map Product / UX / Assurance & compliance evidence

Croissant

MLCommons
open source / Overall score: 2.6

Croissant is MLCommons' metadata format for machine-learning datasets, combining a dataset's metadata, file descriptions, record structure and default ML semantics in one JSON-LD file built on schema.org. The mlcroissant Python library validates and loads Croissant files, and a Responsible AI extension records provenance, collection and bias information. Major dataset hubs publish Croissant descriptions.

Scored as the format and its reference library, not the hubs that serve it. Tagged self-attested: a validator checks the structure, not the truth of what it records.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI
source
public
core features withheld
no

The specification and library are published under Apache-2.0 by MLCommons and developed in the open, with no enterprise, commercial or proprietary directory in the repository. One contributed subdirectory is MIT, which changes nothing.

Adoption

2 medium confidence
2.0

PyPI downloads of mlcroissant, the reference library, measure use of the format's tooling. Hubs that emit Croissant files without the library are not counted.

Capability

3 high confidence
3.0

A Croissant file follows a public schema that anyone can validate with open tooling, so a reader can check its structure without the publisher. Its provenance and responsible-AI fields are still the publisher's own account, two steps below NVIDIA's hardware attestation.

Verified 2026-09-26