论文信息 - Not Quite the Same: Identity Constraints for the Web of Linked Data

Not Quite the Same: Identity Constraints for the Web of Linked Data

Linked Data is based on the idea that information from different sources can flexibly be connected to enable novel applications that individual datasets do not support on their own. This hinges upon the existence of links between datasets that would otherwise be isolated. The most notable form, sameAs links, are intended to express that two identifiers are equivalent in all respects. Unfortunately, many existing ones do not reflect such genuine identity. This study provides a novel method to analyse this phenomenon, based on a thorough theoretical analysis, as well as a novel graph-based method to resolve such issues to some extent. Our experiments on a representative Web-scale set of sameAs links from the Web of Data show that our method can identify and remove hundreds of thousands of constraint violations.

Gerard de Melo

[1] Deborah L. McGuinness,et al. owl:sameAs and Linked Data: An Empirical Study , 2010 .

[2] Gerhard Weikum,et al. Untangling the Cross-Lingual Link Structure of Wikipedia , 2010, ACL.

[3] Deborah L. McGuinness,et al. When owl: sameAs Isn't the Same: An Analysis of Identity in Linked Data , 2010, SEMWEB.

[4] Gerhard Weikum,et al. Language as a Foundation of the Semantic Web , 2008, SEMWEB.

[5] François Scharffe,et al. Final results of the ontology alignment evaluation initiative 2011 , 2011 .

[6] Jürgen Umbrich,et al. Scalable and distributed methods for entity matching, consolidation and disambiguation over linked data corpora , 2012, J. Web Semant..

[7] E. Rosch,et al. Cognition and Categorization , 1980 .

[8] Deborah L. McGuinness,et al. SameAs Networks and Beyond: Analyzing Deployment Status and Implications of owl: sameAs in Linked Data , 2010, International Semantic Web Conference.

[9] L. Barsalou,et al. Ad hoc categories , 1983, Memory & cognition.

[10] Yuval Rabani,et al. ON THE HARDNESS OF APPROXIMATING MULTICUT AND SPARSEST-CUT , 2005, 20th Annual IEEE Conference on Computational Complexity (CCC'05).

[11] Jens Lehmann,et al. DBpedia: A Nucleus for a Web of Open Data , 2007, ISWC/ASWC.