Data model / schema · 2.0 · current

Data Package Standard

Maintained by Data Package Working Group

What it helps you do

Data Package supports a JSON descriptor for a coherent collection of resources, plus Data Resource, Table Schema, and Table Dialect specifications.

  • Discovery
  • Laboratory
  • Cross-cutting
PlanAcquireHarmonizeExchangeLearn + reuse

01

Where it fits and where it does not

Use these four checks before committing implementation time.

Use it when
Lightweight packaging and validation of assay exports, reference tables, tabular analysis results, and other file-based data products.
Limits
It does not supply biological semantics, full provenance, privacy policy, repository trust, or ML-specific intended-use and bias documentation.
Best for
Discovery and Laboratory and Cross-cutting teams working across Harmonize → Exchange → Learn + reuse.
Maturity
ScalingUsable now, but adoption or tooling is still developing. Pilot the exact stack first.

02

See it in the workflow

This view shows the input, the change the standard introduces, and the resulting output.

  1. InputWhat starts

    Discovery and Laboratory and Cross-cutting source data, metadata, and local mappings

  2. Data PackageWhat changes

    Use Data Package as a pinned data model / schema across Harmonize → Exchange → Learn + reuse

  3. OutputWhat becomes possible

    A handoff the next system or team can validate against the same release

Readiness gateIt does not supply biological semantics, full provenance, privacy policy, repository trust, or ML-specific intended-use and bias documentation.

03

A concrete example

A release includes a v2 datapackage.json with stable identity, version, license, sources, contributors, resources, and Table Schemas for tabular files.

Why it matters: Machine-readable resources, field types, missing-value conventions, and categories improve loading and validation, but labels, splits, cohort meaning, and fitness remain external.

04

What it fits with

Can carry DataCite-aligned contributor roles and identifiers, sit inside an RO-Crate, and provide structural metadata beneath a Croissant ML description.

05

Implementation starter

Start with one bounded handoff. Pin, test, and review it before scaling.

  1. Define one handoff, its accountable owner, and the decision Data Package must support.

  2. Pin the exact version and companion artifacts: 2.0 · current.

  3. Map one representative input to the required data model / schema artifacts.

  4. Test the result against the canonical source and record every exception.

  5. Preserve the source data, mappings, and review evidence before scaling.

06

Test the main limitation

Risk

It does not supply biological semantics, full provenance, privacy policy, repository trust, or ML-specific intended-use and bias documentation.

Test

Run one representative end-to-end pilot and record exactly where Data Package loses context, needs an extension, or depends on another standard.

Risk

Machine-readable output may still be unfit for analysis or ML.

Test

Test the output for missing context, provenance, terminology alignment, time leakage, and the intended downstream decision. Machine-readable resources, field types, missing-value conventions, and categories improve loading and validation, but labels, splits, cohort meaning, and fitness remain external.

07

Official resources

Specifications, diagrams, examples, and guides from the organizations that maintain them.

  • Primary source2.0 · current

    Data Package Standard v2

    Official publisher or steward guidance for this data model / schema profile.

    Publisher
    Data Package Working Group
    Open official source

Next action

Put this profile in context

Compare its role with adjacent standards or place it inside an end-to-end data pathway before choosing an implementation.