Data model / schema · 2.0 · current maintained version

GA4GH Phenopackets

Maintained by GA4GH

What it helps you do

Phenopackets supports human- and machine-readable case-level phenotypic, clinical, diagnosis, measurement, biosample, and genomic interpretation data.

  • Clinical
  • Omics
  • Rare disease
PlanAcquireHarmonizeExchangeLearn + reuse

01

Where it fits and where it does not

Use these four checks before committing implementation time.

Use it when
Portable phenotype/genotype exchange for rare disease, cancer, registries, diagnostics, and computational analysis.
Limits
Flexible optionality and ontology dependence require application-specific validation; it does not replace an EHR API, consent layer, or cohort warehouse.
Best for
Clinical and Omics and Rare disease teams working across Harmonize → Exchange → Learn + reuse.
Maturity
ScalingUsable now, but adoption or tooling is still developing. Pilot the exact stack first.

02

See it in the workflow

This view shows the input, the change the standard introduces, and the resulting output.

  1. InputWhat starts

    Clinical and Omics and Rare disease source data, metadata, and local mappings

  2. PhenopacketsWhat changes

    Use Phenopackets as a pinned data model / schema across Harmonize → Exchange → Learn + reuse

  3. OutputWhat becomes possible

    A handoff the next system or team can validate against the same release

Readiness gateFlexible optionality and ontology dependence require application-specific validation; it does not replace an EHR API, consent layer, or cohort warehouse.

03

A concrete example

A rare-disease program validates a v2 Phenopacket with pinned ontology versions, measurements, disease course, biosamples, and variant interpretations.

Why it matters: Provides computable phenotype features and temporal context, while cohort construction, missingness, bias, and leakage controls remain external.

04

What it fits with

Works with GA4GH variation standards and community ontologies, and is designed to interoperate with FHIR; it is not a longitudinal warehouse.

05

Implementation starter

Start with one bounded handoff. Pin, test, and review it before scaling.

  1. Define one handoff, its accountable owner, and the decision Phenopackets must support.

  2. Pin the exact version and companion artifacts: 2.0 · current maintained version.

  3. Map one representative input to the required data model / schema artifacts.

  4. Test the result against the canonical source and record every exception.

  5. Preserve the source data, mappings, and review evidence before scaling.

06

Test the main limitation

Risk

Flexible optionality and ontology dependence require application-specific validation; it does not replace an EHR API, consent layer, or cohort warehouse.

Test

Run one representative end-to-end pilot and record exactly where Phenopackets loses context, needs an extension, or depends on another standard.

Risk

Machine-readable output may still be unfit for analysis or ML.

Test

Test the output for missing context, provenance, terminology alignment, time leakage, and the intended downstream decision. Provides computable phenotype features and temporal context, while cohort construction, missingness, bias, and leakage controls remain external.

07

Official resources

Specifications, diagrams, examples, and guides from the organizations that maintain them.

  • Primary source2.0 · current maintained version

    GA4GH Phenopackets

    Official publisher or steward guidance for this data model / schema profile.

    Publisher
    GA4GH
    Open official source

Next action

Put this profile in context

Compare its role with adjacent standards or place it inside an end-to-end data pathway before choosing an implementation.