Data model / schema · Living specification · documentation 1.50.0

OpenLineage

Maintained by OpenLineage project

What it helps you do

OpenLineage supports runtime and design-time events for jobs, runs, and datasets, with extensible facets for source code, schemas, versions, quality metrics, assertions, and other lineage context.

  • AI / ML
  • Cross-cutting
PlanAcquireHarmonizeExchangeLearn + reuse

01

Where it fits and where it does not

Use these four checks before committing implementation time.

Use it when
Operational lineage instrumentation across ETL, ELT, laboratory, feature-engineering, and model-data pipelines.
Limits
Lineage is only as complete as its instrumentation; inconsistent naming, missing events, facet-version drift, and backend retention can leave an incomplete history.
Best for
AI / ML and Cross-cutting teams working across Acquire → Harmonize → Exchange → Learn + reuse.
Maturity
ScalingUsable now, but adoption or tooling is still developing. Pilot the exact stack first.

02

See it in the workflow

This view shows the input, the change the standard introduces, and the resulting output.

  1. InputWhat starts

    AI / ML and Cross-cutting source data, metadata, and local mappings

  2. OpenLineageWhat changes

    Use OpenLineage as a pinned data model / schema across Acquire → Harmonize → Exchange → Learn + reuse

  3. OutputWhat becomes possible

    A handoff the next system or team can validate against the same release

Readiness gateLineage is only as complete as its instrumentation; inconsistent naming, missing events, facet-version drift, and backend retention can leave an incomplete history.

03

A concrete example

An orchestrated transformation emits stable job, run, and dataset identities plus source revision, input and output versions, schemas, quality assertions, and lifecycle events to a lineage backend.

Why it matters: Supports dataset and feature traceability and quality evidence, but does not establish scientific semantics, consent, label validity, or model fitness.

04

What it fits with

Provides event capture that can feed a PROV-O graph; snapshots and summaries can be linked from RO-Crate, DataCite, DQV, or catalog metadata.

05

Implementation starter

Start with one bounded handoff. Pin, test, and review it before scaling.

  1. Define one handoff, its accountable owner, and the decision OpenLineage must support.

  2. Pin the exact version and companion artifacts: Living specification · documentation 1.50.0.

  3. Map one representative input to the required data model / schema artifacts.

  4. Test the result against the canonical source and record every exception.

  5. Preserve the source data, mappings, and review evidence before scaling.

06

Test the main limitation

Risk

Lineage is only as complete as its instrumentation; inconsistent naming, missing events, facet-version drift, and backend retention can leave an incomplete history.

Test

Run one representative end-to-end pilot and record exactly where OpenLineage loses context, needs an extension, or depends on another standard.

Risk

Machine-readable output may still be unfit for analysis or ML.

Test

Test the output for missing context, provenance, terminology alignment, time leakage, and the intended downstream decision. Supports dataset and feature traceability and quality evidence, but does not establish scientific semantics, consent, label validity, or model fitness.

07

Official resources

Specifications, diagrams, examples, and guides from the organizations that maintain them.

  • Primary sourceLiving specification · documentation 1.50.0

    OpenLineage 1.50.0 object model

    Official publisher or steward guidance for this data model / schema profile.

    Publisher
    OpenLineage project
    Open official source

Next action

Put this profile in context

Compare its role with adjacent standards or place it inside an end-to-end data pathway before choosing an implementation.