论文信息 - Automatic Construction of a Semantic Knowledge Base from CEUR Workshop Proceedings

Automatic Construction of a Semantic Knowledge Base from CEUR Workshop Proceedings

We present an automatic workflow that performs text segmentation and entity extraction from scientific literature to primarily address Task 2 of the Semantic Publishing Challenge 2015. The goal of Task 2 is to extract various information from full-text papers to represent the context in which a document is written, such as the affiliation of its authors and the corresponding funding bodies. Our proposed solution is composed of two subsystems: (i) A text mining pipeline, developed based on the GATE framework, which extracts structural and semantic entities, such as authors’ information and references, and produces semantic (typed) annotations; and (ii) a flexible exporting module, the LODeXporter, which translates the document annotations into RDF triples according to custom mapping rules. Additionally, we leverage existing Named Entity Recognition (NER) tools to extract named entities from text and ground them to their corresponding resources on the Linked Open Data cloud, thus, briefly covering Task 3 objectives, which involves linking of detected entities to resources in existing open datasets. The output of our system is an RDF graph stored in a scalable TDB-based storage with a public SPARQL endpoint for the task’s queries.

Bahar Sateli | René Witte | R. Witte | Bahar Sateli

[1] Kalina Bontcheva,et al. Text Processing with GATE , 2011 .

[2] Bahar Sateli,et al. Supporting Researchers with a Semantic Literature Management Wiki , 2014, SePublica.

[3] Bahar Sateli,et al. What's in this paper?: Combining Rhetorical Entities with Linked Open Data for Semantic Literature Querying , 2015, WWW.

[4] Siegfried Handschuh,et al. SALT - Semantically Annotated LaTeX for scientific publications , 2007 .

[5] Fabio Vitali,et al. The Document Components Ontology (DoCO) , 2016, Semantic Web.