- Helps with
- Findability, accessibility, interoperability, and reusability of data, metadata, and infrastructure.
- Best for
- Use as the outcome framework and assessment lens across the full R&D lifecycle.
- Watch out
- Principles describe desired behavior, not a single technical architecture or conformance test.
Profile library
Browse the knowledge base
Search by task, domain, or lifecycle stage, then check each standard's fit and limits.
35 of 93 profiles
35 of 93 profiles · Select up to three to compare.
- Helps with
- Protocol-to-analysis clinical and nonclinical research data, including acquisition, tabulation, analysis, and submission structures.
- Best for
- Best for regulated studies and traceable submission packages; less natural for early discovery or raw instrument output.
- Watch out
- Conformance is detailed and version-sensitive; transformations can preserve structure while losing source context.
- Helps with
- Experimental design, sample characteristics, protocols, assay technologies, and sample-to-data relationships.
- Best for
- Multi-omics and multi-assay study metadata at the boundary between experiment design and repository submission.
- Watch out
- Rich metadata entry is labor-intensive and local templates can drift without governance and validation.
- Helps with
- A family of interoperable biological and biomedical ontologies governed by principles for openness, scope, identifiers, relations, and maintenance.
- Best for
- Semantic annotation, knowledge graphs, terminology normalization, and cross-dataset integration.
- Watch out
- Coverage and maintenance vary by ontology; overlap, versioning, and term-selection policy still require local stewardship.
- Helps with
- Persistent and resolvable identifiers, predictable metadata retrieval, and typing for machine-actionable digital objects.
- Best for
- Long-horizon infrastructure design for object-level interoperability across repositories and automated agents.
- Watch out
- The available documentation is explicitly incomplete and should not be treated as a final specification.
- Helps with
- Standards-agnostic biomedical concept definitions plus implementation artifacts such as SDTM Dataset Specializations and value-level metadata.
- Best for
- Computable clinical concepts that connect protocol, collection design, tabulation metadata, and downstream automation.
- Watch out
- Content is informative and incrementally curated; concept, specialization, terminology, and downstream standard versions must be pinned together.
- Helps with
- Identification of medicinal products, pharmaceutical products, substances, dose forms, routes, units, and packages across the product lifecycle.
- Best for
- Regulated medicinal-product master data, cross-system product identity, and jurisdictional submissions such as EMA SPOR/PMS.
- Watch out
- The suite spans multiple ISO standards and amendments; implementation scope, identifiers, and timelines vary by jurisdiction and are still evolving.
- Helps with
- Standardized biomedical data-use permission terms for matching controlled-access datasets to research purposes.
- Best for
- Consent-aware discovery, data access review, and machine-readable permitted-use conditions in genomics and health research.
- Watch out
- Ontology matching cannot resolve jurisdiction, contract, consent nuance, expiry, or downstream duties without authoritative policy and human governance.
- Helps with
- Govern, Map, Measure, and Manage functions for addressing AI risks across organizations and system lifecycles.
- Best for
- Governance overlay for intended use, accountability, risk measurement, release decisions, and ongoing monitoring.
- Watch out
- Voluntary and use-case agnostic; it does not prescribe life-science schemas, legal compliance, or quantitative acceptance thresholds.
- Helps with
- Tabular sample-to-data relationships, biological and technical factors, replicates, instruments, acquisition context, and proteomics experimental design.
- Best for
- Proteomics studies that need an explicit, machine-readable mapping from biosamples and factors to raw and processed mass-spectrometry files.
- Watch out
- It does not encode downstream statistical-analysis parameters or results. Working-branch templates and rules can move ahead of the final PSI specification, so release and validator versions must be pinned.
- Helps with
- Minimum information for interpretable and reproducible microarray and high-throughput sequencing studies, including design, samples, raw and processed data, and protocols.
- Best for
- A submission and publication completeness gate for functional-genomics studies, especially when preparing repository records and supporting data.
- Watch out
- These are minimum-information checklists rather than one executable schema. Stewardship is legacy, and newer assay classes such as single-cell data require current repository guidance and additional profiles.
- Helps with
- Reusable indicators, priorities, maturity levels, and evaluation guidance for assessing data and metadata against the FAIR principles.
- Best for
- Use as the evidence rubric that turns FAIR from an aspiration into a repeatable release and improvement assessment.
- Watch out
- It is not a certification, and locally adapted scoring or weighting means totals from different assessment tools are not automatically comparable.
- Helps with
- Machine-readable concepts for data and processing, purposes, legal bases, parties, recipients, rights, risks, controls, technologies, AI, and jurisdiction-specific laws.
- Best for
- Privacy and data-protection context for human, genomic, clinical, real-world, and AI datasets where DUO alone is too narrow.
- Watch out
- DPV is a vocabulary, not legal advice or an enforcement engine; local authority, consent, contracts, and jurisdiction-specific interpretation remain controlling.
- Helps with
- Policies containing permissions, prohibitions, duties, parties, assets, constraints, inheritance, and conflict strategies.
- Best for
- Machine-readable usage conditions for datasets and distributions, including research-purpose, redistribution, attribution, retention, and temporal or jurisdictional constraints.
- Watch out
- A syntactically valid policy does not prove the assigner has authority, make the policy legally enforceable, or provide the system that evaluates and enforces it.
ISO/IEC 5259 Data Quality for Analytics and Machine Learning
- Helps with
- Terminology and examples, data-quality measures, management requirements, a process framework, and a governance framework for analytics and ML data.
- Best for
- The quality-management spine for training, validation, and evaluation data used in life-science analytics and ML.
- Watch out
- The normative publications are not freely available, the series is cross-domain, and it does not provide life-science thresholds, domain semantics, or regulatory approval.
- Helps with
- Structured, audience-aware summaries of dataset origins, collection and annotation, intended use, evaluation context, ethical considerations, and decisions affecting downstream performance.
- Best for
- Human-facing readiness and release documentation for clinical, imaging, omics, laboratory, and real-world ML datasets.
- Watch out
- There is no single mandatory schema or conformance test; narrative claims require linked evidence, ownership, review, and update controls.
- Helps with
- A computable study definition spanning objectives, endpoints, eligibility, interventions, schedule of activities, amendments, estimands, and protocol content.
- Best for
- Upstream protocol facts that must flow consistently into study-build, registry, document, and downstream data systems.
- Watch out
- USDM is a model and reference architecture. It is not an EDC, submission dataset, or proof that generated documents comply with every regulator. Every linked terminology and API version must be pinned.
- Helps with
- Portable JSON or YAML descriptions of command-line tools and data-intensive workflows, including typed inputs and outputs, requirements, dependencies, conditional steps, and scatter execution.
- Best for
- Reproducible bioinformatics and scientific workflows that must move across workstations, clusters, clouds, and compatible workflow engines.
- Watch out
- Runner behavior outside the specified execution model can vary, and CWL alone does not freeze containers, reference data, credentials, resource policies, or scientific assumptions.
- Helps with
- Contextual metadata about the source, environment, sampling, processing, and sequencing of genomes, metagenomes, marker genes, single amplified genomes, metagenome-assembled genomes, and uncultivated virus genomes.
- Best for
- Sequence and microbiome studies that need repository-ready sample context through a named checklist plus an environmental or host extension.
- Watch out
- The tagged release and the maintained main schema can diverge, while repository profiles may lag or alter requirements. Pin the schema commit, checklist, extension, term identifiers, and target repository profile.
- Helps with
- Structured communication of high-throughput sequencing analyses through provenance, usability, description, execution, parameters, inputs and outputs, error domains, attribution, and review metadata.
- Best for
- Bioinformatics analyses that need a stable, reviewable account for scientific exchange, regulated communication, or reproducibility assessment.
- Watch out
- BioCompute is descriptive rather than executable and does not prove analytical validity or scientific fitness. The IEEE edition, open schema, object version, and referenced workflow assets must be pinned together.
Recommended Metadata for Biological Images
- Helps with
- Study, study-component, biosample, specimen, image-acquisition, image-data, image-correlation, analysis, and annotation context for reusable biological images.
- Best for
- Biological imaging studies that need scientific context around raw images, processed images, derived annotations, and repository deposits.
- Watch out
- Version 1.5 is the BioImage Archive operational model and should not be presented as a universal canonical serialization. REMBI guidance does not provide one cross-repository conformance test.
- Helps with
- The Essential 10 and Recommended Set for study design, sample size, inclusion and exclusion, randomization, blinding, outcome measures, statistics, animal characteristics, procedures, results, and interpretation.
- Best for
- Planning, recording, reporting, and reviewing in vivo experiments so readers can assess methodological rigor and reproduce the work.
- Watch out
- ARRIVE is a human-facing reporting guideline rather than a machine-readable study schema. Checklist completion does not establish ethical approval, statistical validity, or reproducibility.
RDA DMP Common Standard for Machine-actionable Data Management Plans
- Helps with
- A machine-actionable application profile for projects, datasets, distributions, contributors, funding, costs, repositories, licenses, identifiers, security, privacy, quality assurance, preservation, and technical resources across a data-management plan.
- Best for
- Planning and maintaining computable data-governance commitments that can move among researchers, funders, repositories, and infrastructure providers.
- Watch out
- The 2019 RDA endorsement and the later v1.2 repository release are distinct status claims; the profile does not impose funder rules, prove that declared actions occurred, or replace domain metadata.
- Helps with
- Frameworks and metamodels for registering metadata items with governed identification, naming, definitions, classification, mappings, registration status, administration, and dataset descriptions.
- Best for
- Enterprise or federated registries that must govern the meaning and lifecycle of data elements, concepts, value domains, classifications, and related metadata independently of a storage platform.
- Watch out
- The suite is an abstract metamodel rather than a ready-made catalog API; parts version independently, much normative text is not freely available, and useful deployment requires a scoped profile and governance process.
- Helps with
- A YAML data contract spanning fundamentals, schema, references, data-quality rules, support channels, pricing, teams, roles, service-level agreements, infrastructure, and extension properties.
- Best for
- A versioned producer-consumer agreement for operational datasets, tables, streams, or APIs where schema, ownership, quality expectations, and service levels must be testable together.
- Watch out
- ODCS is an evolving open industry standard rather than an ISO or W3C standard; its companion JSON Schema does not supersede the specification, and declared rules do not prove scientific fitness.
ISO/IEC TS 27560 Consent Record Information Structure
- Helps with
- An interoperable, open, and extensible information structure for consent records and receipts, including exchange and lifecycle management of consent for processing personally identifiable information.
- Best for
- Systems that must record what a person consented to, provide a receipt, exchange consent information, and manage changes or withdrawal over time.
- Watch out
- A conforming record does not prove that consent was informed, freely given, current, or the correct lawful basis; jurisdictional and ethical requirements remain controlling, and the standard is being revised.
ISO 14721 Open Archival Information System Reference Model
- Helps with
- Archive responsibilities, designated communities, ingest, archival storage, data management, preservation planning, administration, access, and submission, archival, and dissemination information packages.
- Best for
- Designing or evaluating repositories that accept responsibility for preserving information and keeping it understandable and available to a defined community over technological and organizational change.
- Watch out
- OAIS is a conceptual reference model, not software or certification; the term Open refers to standards development rather than unrestricted data access, and local preservation policies remain necessary.
CARE Principles for Indigenous Data Governance
- Helps with
- Collective Benefit, Authority to Control, Responsibility, and Ethics for the governance and use of Indigenous data and Indigenous Knowledge.
- Best for
- Any data activity involving Indigenous peoples, lands, resources, knowledge, or community interests where openness and reuse must be governed through rights, authority, benefit, and accountability.
- Watch out
- CARE is not a checklist that an external organization can self-certify; obligations are community and context specific, and implementation requires continuing participation by legitimate Indigenous authorities.
ICH E6(R3) Guideline for Good Clinical Practice
- Helps with
- Participant protection, reliable trial results, proportional quality management, data governance, source records and metadata, audit trails, computerized systems, oversight, and essential records.
- Best for
- Any regulated interventional clinical trial, including decentralized, pragmatic, or RWD-enabled designs that need an accountable quality and data-integrity operating model.
- Watch out
- E6(R3) is not a data schema or executable validator, and Step 4 publication does not mean identical implementation dates or legal requirements in every region.
ICH M14 General Principles on Planning, Designing, Analysing, and Reporting of Non-interventional Studies That Utilise Real-World Data for Safety Assessment of Medicines
- Helps with
- Planning, data-source evaluation, study design, variable validation, bias and confounding assessment, analysis, sensitivity work, and reporting for non-interventional medicine-safety studies using RWD.
- Best for
- Regulatory pharmacoepidemiology and post-authorization safety studies where routinely collected data must support an explicit, reviewable evidence argument.
- Watch out
- The guideline is focused on non-interventional safety assessment, not every effectiveness question, and conformance cannot make an inadequate source complete or remove unmeasured confounding.
CDISC Trial Master File Reference Model
- Helps with
- A standardized taxonomy, artifact identifiers, zones, sections, levels, metadata, and filing conventions for essential clinical-trial records in sponsor and investigator files.
- Best for
- Trial record management, sponsor-CRO transfer, completeness assessment, inspection readiness, and reconstruction of trial conduct across organizations and systems.
- Watch out
- CDISC describes the model as adaptable rather than a formal enforceable standard. Company SOPs, jurisdictional requirements, and the upcoming successor model can change artifact selection and filing rules.
- Helps with
- Aggregate metadata for biobanks, sample and data collections, research resources, and networks, with separate individual-level components for samples, donors, and events.
- Best for
- Biobank directories, cohort discovery, federated catalogues, access negotiation, and metadata exchange across biospecimen networks.
- Watch out
- Core 3.0 is aggregate-level metadata, not a complete specimen record. Individual-level components and serializations version separately, and MIABIS does not by itself establish consent, sample quality, or one universal conformance test.
- Helps with
- Requirements for biobank competence, impartiality, consistent operation, and quality control of biological material and associated data collections.
- Best for
- Biobank quality systems, capability assessment, accreditation preparation, partner qualification, and governance of materials intended for research and development.
- Watch out
- The current edition is under revision, so contracts and assessments must pin the exact edition. Its scope excludes biological material intended for food, feed, or therapeutic use.
- Helps with
- Adaptive immune receptor repertoire study metadata, sample processing, raw-sequence references, rearrangements, repertoires, germline and genotype records, clones, single-cell data, reactivity, and receptor annotations.
- Best for
- AIRR-seq publication, repository submission, immune-repertoire exchange, federated discovery, reproducible analysis, and cross-tool interoperability.
- Watch out
- MiAIRR compliance establishes reporting completeness, not sequencing quality or biological validity. The v2.0 transition, optional fields, ontology versions, germline references, privacy, and repository profiles require separate controls.
GA4GH Experiments Metadata Checklist
- Helps with
- Minimum experiment-level properties for high-throughput sequencing, focused on library preparation, sequencing instruments, run context, protocols, and relevant identifiers.
- Best for
- Repositories, laboratories, and federated discovery systems that need comparable descriptions of WGS, RNA-seq, methylation, and other sequencing experiments.
- Watch out
- The current scope excludes biological sample descriptors, clinical data, downstream processing, and analysis. A completed checklist is not a reproducibility record, and the checklist does not require one canonical serialization.