This step establishes which controlled vocabularies, terminologies and ontologies will be used to describe the research object and its associated metadata, and applies their terms in ways that support consistent interpretation and reuse.
Semantic resources provide agreed identifiers, labels, definitions and relationships for concepts. They can make the meaning of research object content and metadata clearer to people and machines and help different research objects, services and communities refer consistently to the same concepts. A controlled vocabulary may define a permitted list of values, while an ontology may also provide formal relationships and constraints. The FAIRification activity should select the level of semantic detail needed for its intended uses. Use this step during project examination to identify semantic requirements, inventory current vocabulary use, assess available capabilities and select candidate resources. During an implementation cycle, use it to extend or develop semantic resources where necessary, apply annotations and establish the required management arrangements.
Identify research object types
Select or define a domain model
Identifier minting
Identifier discovery and reuse
Vocabulary discovery and selection
Vocabulary extension and development
Semantic annotation
Vocabulary management
Vocabulary discovery and selection
Find, assess and choose semantic resources that provide appropriate identifiers and descriptions for the concepts represented in the research object and its associated metadata.
This capability includes determining which concepts need controlled terms, identifying relevant community resources and evaluating whether those resources provide sufficient coverage, granularity, semantic precision and operational support.
The capability may be provided through project expertise, domain communities, vocabulary registries, lookup services, ontology portals, standards organisations, repositories or other FAIR-enabling resources.
Resources
3 ELIXIR Stories items
- Plant Sciences Community Showcase Cascade mapping from FT 2–6: integration of multiple resources, data harmonisation
- Mapping the DCAT-based EJP RD Metadata Model to Bioschemas for improved interoperability Cascade mapping from FT 5: choose data vocabularies: selecting, developing and annotating with data vocabularies: use of DCAT, DCT, FOAF, EJP RD
- Galaxy: a story about interoperability Cascade mapping from FT 5: Data vocabularies, 6. Transform data for interop: all tools rendered interoperable by describing I/O and filetypes with EDAM, cross referenced with registries
3 FAIR Metroline items
- Use ontologies in the model Choosing ontology terms that make concepts explicit and machine readable is a key part of selecting vocabularies.
- Create or reuse a semantic (meta)data model The concepts and relations defined in a semantic model guide which vocabularies should be selected.
- Creating a FAIR Implementation Profile (FIP) Recording chosen semantic resources in a FIP makes vocabulary selection explicit at community level.
2 FAIR Cookbook items
Vocabulary extension and development
Add or develop new concepts when established semantic resources do not adequately cover the requirements of the FAIRification activity.
This capability may involve requesting a new term from an existing authority, contributing corrections or relationships, creating a governed local extension, defining an application ontology or developing a new semantic resource.
The preferred sequence is to:
- Reuse a suitable existing term.
- Request a new term, definition or correction from the maintaining authority.
- Use an established extension mechanism.
- Create a governed local extension that reuses existing identifiers where possible.
- Combine or extract modules from compatible semantic resources.
- Develop a new vocabulary or ontology only where no suitable alternative exists.
Vocabulary development usually requires coordination with domain communities, intended users, ontology or terminology specialists, repositories and software implementers. The required capability may therefore be provided outside the immediate FAIRification team.
Resources
3 ELIXIR Stories items
- Plant Sciences Community Showcase Cascade mapping from FT 2–6: integration of multiple resources, data harmonisation
- Mapping the DCAT-based EJP RD Metadata Model to Bioschemas for improved interoperability Cascade mapping from FT 5: choose data vocabularies: selecting, developing and annotating with data vocabularies: use of DCAT, DCT, FOAF, EJP RD
- Galaxy: a story about interoperability Cascade mapping from FT 5: Data vocabularies, 6. Transform data for interop: all tools rendered interoperable by describing I/O and filetypes with EDAM, cross referenced with registries
2 FAIR Metroline items
- Create or reuse a semantic (meta)data model Defining the concepts needed in the model creates the foundation for developing vocabularies.
- Use ontologies in the model Newly developed terms can become part of the ontology driven model.
1 FAIR Cookbook items
Semantic annotation
Associate research objects, components, metadata elements and values with identifiable concepts from selected semantic resources.
This capability includes selecting the correct concept, representing its identifier in the appropriate context and recording sufficient provenance to understand how and why the annotation was made.
In addition to a term from a vocabulary, an annotation usually also includes the subject being described and the relationship between the subject and the concept that the term represents. For example, stating that a research object “is about” a disease, “uses” a method or “has specimen type” a biological material expresses different meanings.
The capability may be provided through manual curation, data-entry systems, transformation workflows, text-mining or annotation tools, repository services or combinations of automated and expert processes.
Resources
3 ELIXIR Stories items
- Plant Sciences Community Showcase Cascade mapping from FT 2–6: integration of multiple resources, data harmonisation
- Mapping the DCAT-based EJP RD Metadata Model to Bioschemas for improved interoperability Cascade mapping from FT 5: choose data vocabularies: selecting, developing and annotating with data vocabularies: use of DCAT, DCT, FOAF, EJP RD
- Galaxy: a story about interoperability Cascade mapping from FT 5: Data vocabularies, 6. Transform data for interop: all tools rendered interoperable by describing I/O and filetypes with EDAM, cross referenced with registries
2 FAIR Metroline items
- Use ontologies in the model Linking data and metadata to formal ontology or vocabulary terms gives annotation its meaning.
- Apply (meta)data model Putting the chosen semantic model into practice includes annotating data and metadata with the right terms.
2 FAIR Cookbook items
Vocabulary management
Maintain reliable and reproducible use of semantic resources over time.
For most FAIRification activities, this means managing the project’s use of externally maintained vocabularies: recording versions, monitoring changes, updating annotations and preserving reproducibility.
Where the project or community maintains a vocabulary, ontology, value set or extension, the capability additionally includes editorial governance, identifiers, releases, publication, support and long-term maintenance.
Resources
3 ELIXIR Stories items
- Plant Sciences Community Showcase Cascade mapping from FT 2–6: integration of multiple resources, data harmonisation
- Mapping the DCAT-based EJP RD Metadata Model to Bioschemas for improved interoperability Cascade mapping from FT 5: choose data vocabularies: selecting, developing and annotating with data vocabularies: use of DCAT, DCT, FOAF, EJP RD
- Galaxy: a story about interoperability Cascade mapping from FT 5: Data vocabularies, 6. Transform data for interop: all tools rendered interoperable by describing I/O and filetypes with EDAM, cross referenced with registries
2 FAIR Metroline items
- Creating a FAIR Implementation Profile (FIP) Documenting and maintaining community choices over time supports vocabulary management.
- Use ontologies in the model Sustained and consistent use of ontologies depends on active vocabulary management.