Standard · Sequences v2.0 · Sequence Collections v1.0.0

GA4GH refget Sequences and Sequence Collections

Maintained by GA4GH Federated Analysis Work Stream

What it helps you do

refget + SeqCol supports content-derived identifiers and retrieval for individual reference sequences, plus identification, lookup, and comparison of sequence collections such as genomes, transcriptomes, and proteomes.

  • Omics
  • Bioinformatics
PlanAcquireHarmonizeExchangeLearn + reuse

01

Where it fits and where it does not

Use these four checks before committing implementation time.

Use it when
Genomic pipelines, archives, variant systems, and federated services that must prove which reference sequence or assembly collection an analysis used.
Limits
Content identity does not establish that a reference is authoritative, biologically appropriate, complete, or suitable for a particular analysis. Sequence Collections is newer, so service compatibility and collection conventions require explicit testing.
Best for
Omics and Bioinformatics teams working across Acquire → Harmonize → Exchange → Learn + reuse.
Maturity
ScalingUsable now, but adoption or tooling is still developing. Pilot the exact stack first.

02

See it in the workflow

This view shows the input, the change the standard introduces, and the resulting output.

  1. InputWhat starts

    Omics and Bioinformatics source data, metadata, and local mappings

  2. refget + SeqColWhat changes

    Use refget + SeqCol as a pinned standard across Acquire → Harmonize → Exchange → Learn + reuse

  3. OutputWhat becomes possible

    A handoff the next system or team can validate against the same release

Readiness gateContent identity does not establish that a reference is authoritative, biologically appropriate, complete, or suitable for a particular analysis. Sequence Collections is newer, so service compatibility and collection conventions require explicit testing.

03

A concrete example

Compute and persist refget digests for sequences, generate the applicable Sequence Collection identifiers, retain aliases and collection metadata, and verify that retrieved content reproduces the recorded identifiers.

Why it matters: Creates reproducible reference keys for feature joins and provenance, while biological assembly choice, coordinate liftover, and reference bias remain separate analytical decisions.

04

What it fits with

Provides reference identity for HGVS, GA4GH VRS, HTS formats, htsget, and other coordinate-based genomic artifacts; aliases and assembly labels remain complementary metadata.

05

Implementation starter

Start with one bounded handoff. Pin, test, and review it before scaling.

  1. Define one handoff, its accountable owner, and the decision refget + SeqCol must support.

  2. Pin the exact version and companion artifacts: Sequences v2.0 · Sequence Collections v1.0.0.

  3. Map one representative input to the required standard artifacts.

  4. Test the result against the canonical source and record every exception.

  5. Preserve the source data, mappings, and review evidence before scaling.

06

Test the main limitation

Risk

Content identity does not establish that a reference is authoritative, biologically appropriate, complete, or suitable for a particular analysis. Sequence Collections is newer, so service compatibility and collection conventions require explicit testing.

Test

Run one representative end-to-end pilot and record exactly where refget + SeqCol loses context, needs an extension, or depends on another standard.

Risk

Machine-readable output may still be unfit for analysis or ML.

Test

Test the output for missing context, provenance, terminology alignment, time leakage, and the intended downstream decision. Creates reproducible reference keys for feature joins and provenance, while biological assembly choice, coordinate liftover, and reference bias remain separate analytical decisions.

07

Official resources

Specifications, diagrams, examples, and guides from the organizations that maintain them.

  • Primary sourcev2.0

    refget Sequences

    The GA4GH product page for content-derived identifiers and retrieval of individual reference sequences.

    Publisher
    GA4GH
    Open official source
  • Specificationv1.0.0

    refget Sequence Collections

    The companion GA4GH product for identifying, comparing, and looking up collections such as genomes, transcriptomes, and proteomes.

    Publisher
    GA4GH
    Open official source

Next action

Put this profile in context

Compare its role with adjacent standards or place it inside an end-to-end data pathway before choosing an implementation.