Framework · 2022 framework + playbook

Data Cards for AI Dataset Documentation

Maintained by Google Research Data Cards authors

What it helps you do

Data Cards supports structured, audience-aware summaries of dataset origins, collection and annotation, intended use, evaluation context, ethical considerations, and decisions affecting downstream performance.

  • AI / ML
  • Cross-cutting
PlanAcquireHarmonizeExchangeLearn + reuse

01

Where it fits and where it does not

Use these four checks before committing implementation time.

Use it when
Human-facing readiness and release documentation for clinical, imaging, omics, laboratory, and real-world ML datasets.
Limits
There is no single mandatory schema or conformance test; narrative claims require linked evidence, ownership, review, and update controls.
Best for
AI / ML and Cross-cutting teams working across Plan → Acquire → Harmonize → Exchange → Learn + reuse.
Maturity
ScalingUsable now, but adoption or tooling is still developing. Pilot the exact stack first.

02

See it in the workflow

This view shows the input, the change the standard introduces, and the resulting output.

  1. InputWhat starts

    AI / ML and Cross-cutting source data, metadata, and local mappings

  2. Data CardsWhat changes

    Use Data Cards as a pinned framework across Plan → Acquire → Harmonize → Exchange → Learn + reuse

  3. OutputWhat becomes possible

    A handoff the next system or team can validate against the same release

Readiness gateThere is no single mandatory schema or conformance test; narrative claims require linked evidence, ownership, review, and update controls.

03

A concrete example

A release card documents provenance, sample or cohort construction, annotation and QC, split design, intended and out-of-scope uses, relevant coverage and performance evidence, risks, and version changes.

Why it matters: Makes the rationale and limitations that determine responsible reuse visible to human reviewers, while companion machine-readable metadata is still required for automation.

04

What it fits with

Complements machine-readable Croissant and SPDX records, ISO/IEC 5259 quality evidence, and NIST AI RMF governance; Datasheets for Datasets is a closely related predecessor.

05

Implementation starter

Start with one bounded handoff. Pin, test, and review it before scaling.

  1. Define one handoff, its accountable owner, and the decision Data Cards must support.

  2. Pin the exact version and companion artifacts: 2022 framework + playbook.

  3. Map one representative input to the required framework artifacts.

  4. Test the result against the canonical source and record every exception.

  5. Preserve the source data, mappings, and review evidence before scaling.

06

Test the main limitation

Risk

There is no single mandatory schema or conformance test; narrative claims require linked evidence, ownership, review, and update controls.

Test

Run one representative end-to-end pilot and record exactly where Data Cards loses context, needs an extension, or depends on another standard.

Risk

Machine-readable output may still be unfit for analysis or ML.

Test

Test the output for missing context, provenance, terminology alignment, time leakage, and the intended downstream decision. Makes the rationale and limitations that determine responsible reuse visible to human reviewers, while companion machine-readable metadata is still required for automation.

07

Official resources

Specifications, diagrams, examples, and guides from the organizations that maintain them.

  • Primary source2022 framework + playbook

    Google Research Data Cards paper

    Official publisher or steward guidance for this framework profile.

    Publisher
    Google Research Data Cards authors
    Open official source

Next action

Put this profile in context

Compare its role with adjacent standards or place it inside an end-to-end data pathway before choosing an implementation.