BiobankConnect: software to rapidly connect data elements for pooled analysis across biobanks using ontological and lexical indexing

Chao Pang; Dennis Hendriksen; Martijn Dijkstra; K Joeri van der Velde; Joel Kuiper; Hans L Hillege; Morris A Swertz

doi:10.1136/amiajnl-2013-002577

BiobankConnect: software to rapidly connect data elements for pooled analysis across biobanks using ontological and lexical indexing

J Am Med Inform Assoc. 2015 Jan;22(1):65-75. doi: 10.1136/amiajnl-2013-002577. Epub 2014 Oct 31.

Authors

Chao Pang¹, Dennis Hendriksen², Martijn Dijkstra², K Joeri van der Velde³, Joel Kuiper¹, Hans L Hillege⁴, Morris A Swertz³

Affiliations

¹ Department of Genetics, Genomics Coordination Center, University of Groningen, University Medical Center Groningen, Groningen, The Netherlands Department of Epidemiology, University of Groningen, University Medical Center Groningen, Groningen, The Netherlands.
² Department of Genetics, Genomics Coordination Center, University of Groningen, University Medical Center Groningen, Groningen, The Netherlands.
³ Department of Genetics, Genomics Coordination Center, University of Groningen, University Medical Center Groningen, Groningen, The Netherlands Groningen Bioinformatics Center, University of Groningen, Groningen, The Netherlands.
⁴ Department of Epidemiology, University of Groningen, University Medical Center Groningen, Groningen, The Netherlands.

Abstract

Objective: Pooling data across biobanks is necessary to increase statistical power, reveal more subtle associations, and synergize the value of data sources. However, searching for desired data elements among the thousands of available elements and harmonizing differences in terminology, data collection, and structure, is arduous and time consuming.

Materials and methods: To speed up biobank data pooling we developed BiobankConnect, a system to semi-automatically match desired data elements to available elements by: (1) annotating the desired elements with ontology terms using BioPortal; (2) automatically expanding the query for these elements with synonyms and subclass information using OntoCAT; (3) automatically searching available elements for these expanded terms using Lucene lexical matching; and (4) shortlisting relevant matches sorted by matching score.

Results: We evaluated BiobankConnect using human curated matches from EU-BioSHaRE, searching for 32 desired data elements in 7461 available elements from six biobanks. We found 0.75 precision at rank 1 and 0.74 recall at rank 10 compared to a manually curated set of relevant matches. In addition, best matches chosen by BioSHaRE experts ranked first in 63.0% and in the top 10 in 98.4% of cases, indicating that our system has the potential to significantly reduce manual matching work.

Conclusions: BiobankConnect provides an easy user interface to significantly speed up the biobank harmonization process. It may also prove useful for other forms of biomedical data integration. All the software can be downloaded as a MOLGENIS open source app from http://www.github.com/molgenis, with a demo available at http://www.biobankconnect.org.

Keywords: Biobank; Data integration; Harmonization; Search.

Publication types

Research Support, Non-U.S. Gov't

MeSH terms

Abstracting and Indexing*
Biological Ontologies*
Computational Biology*
Datasets as Topic*
Humans
Software*
Systems Integration
User-Computer Interface