01
Where it fits and where it does not
Use these four checks before committing implementation time.
- Use it when
- Multi-source cohort analytics, patient-level prediction, characterization, and network studies after ETL.
- Limits
- ETL is expensive, source nuance can be compressed, and vocabulary maintenance is an ongoing operational dependency.
- Best for
- Real-world evidence and Clinical teams working across Harmonize → Learn + reuse.
- Maturity
- EstablishedSuitable for production assessment. Pin the exact release and any implementation profile.
02
See it in the workflow
This view shows the input, the change the standard introduces, and the resulting output.
- InputWhat starts
Real-world evidence and Clinical source data, metadata, and local mappings
- OMOP CDMWhat changes
Use OMOP CDM as a pinned data model / schema across Harmonize → Learn + reuse
- OutputWhat becomes possible
A handoff the next system or team can validate against the same release
03
A concrete example
Hospitals transform EHR and claims sources into a common CDM to run a shared phenotype and outcomes protocol.
Why it matters: A consistent longitudinal feature surface is valuable for ML, but label design, missingness, site shift, and temporal leakage remain local responsibilities.
04
What it fits with
Source vocabularies map to OHDSI standard concepts; FHIR and claims data commonly feed ETL; analysis tools consume the CDM.
- TerminologyLOINC
Both support Clinical and Real-world evidence work and meet around Harmonize, Learn + reuse. Compare their roles before treating them as interchangeable.
Explore relationship - TerminologySNOMED CT
Both support Clinical and Real-world evidence work and meet around Harmonize, Learn + reuse. Compare their roles before treating them as interchangeable.
Explore relationship - Data model / schemaPCORnet CDM
Both support Real-world evidence and Clinical work and meet around Harmonize, Learn + reuse. Compare their roles before treating them as interchangeable.
Explore relationship - Governance frameworkICH E6(R3)
Both support Clinical and Real-world evidence work and meet around Harmonize, Learn + reuse. Compare their roles before treating them as interchangeable.
Explore relationship
05
Implementation starter
Start with one bounded handoff. Pin, test, and review it before scaling.
Define one handoff, its accountable owner, and the decision OMOP CDM must support.
Pin the exact version and companion artifacts: CDM 5.4.
Map one representative input to the required data model / schema artifacts.
Test the result against the canonical source and record every exception.
Preserve the source data, mappings, and review evidence before scaling.
06
Test the main limitation
ETL is expensive, source nuance can be compressed, and vocabulary maintenance is an ongoing operational dependency.
Run one representative end-to-end pilot and record exactly where OMOP CDM loses context, needs an extension, or depends on another standard.
Machine-readable output may still be unfit for analysis or ML.
Test the output for missing context, provenance, terminology alignment, time leakage, and the intended downstream decision. A consistent longitudinal feature surface is valuable for ML, but label design, missingness, site shift, and temporal leakage remain local responsibilities.
07
Official resources
Specifications, diagrams, examples, and guides from the organizations that maintain them.
OHDSI OMOP CDM documentation
Official publisher or steward guidance for this data model / schema profile.
- Publisher
- OHDSI