Computing identity co-reference across drug discovery datasets

Christian Y A Brenninkmeijer*, Ian Dunlop, Carole Goble, Alasdair J G Gray, Steve Pettifer, Robert Stevens

*Corresponding author for this work

Research output: Chapter in Book/Report/Conference proceedingConference contribution

1 Citation (Scopus)
81 Downloads (Pure)

Abstract

This paper presents the rules used within the Open PHACTS (http://www.openphacts.org) Identity Management Service to compute co-reference chains across multiple datasets. The web of (linked) data has encouraged a proliferation of identifiers for the concepts captured in datasets; with each dataset using their own identifier. A key data integration challenge is linking the co-referent identifiers, i.e. identifying and linking the equivalent concept in every dataset. Exacerbating this challenge, the datasets model the data difierently, so when is one representation truly the same as another? Finally, different users have their own task and domain specific notions of equivalence that are driven by their operational knowledge. Consumers of the data need to be able to choose the notion of operational equivalence to be applied for the context of their application. We highlight the challenges of automatically computing co-reference and the need for capturing the context of the equivalence. This context is then used to control the co-reference computation. Ultimately, the context will enable data consumers to decide which co-references to include in their applications.

Original languageEnglish
Title of host publicationProceedings of the 6th International Workshop on Semantic Web Applications and Tools for Life Sciences
Volume1114
Publication statusPublished - 2013
Event6th International Workshop on Semantic Web Applications and Tools for Life Sciences - Edinburgh, United Kingdom
Duration: 9 Dec 201312 Dec 2013

Workshop

Workshop6th International Workshop on Semantic Web Applications and Tools for Life Sciences
Country/TerritoryUnited Kingdom
CityEdinburgh
Period9/12/1312/12/13

Keywords

  • Linked Data
  • DATA INTEGRATION
  • Mapping

Fingerprint

Dive into the research topics of 'Computing identity co-reference across drug discovery datasets'. Together they form a unique fingerprint.

Cite this