论文信息 - Predicting speculation: a simple disambiguation approach to hedge detection in biomedical literature

Predicting speculation: a simple disambiguation approach to hedge detection in biomedical literature

BackgroundThis paper presents a novel approach to the problem of hedge detection, which involves identifying so-called hedge cues for labeling sentences as certain or uncertain. This is the classification problem for Task 1 of the CoNLL-2010 Shared Task, which focuses on hedging in the biomedical domain. We here propose to view hedge detection as a simple disambiguation problem, restricted to words that have previously been observed as hedge cues. As the feature space for the classifier is still very large, we also perform experiments with dimensionality reduction using the method of random indexing.ResultsThe SVM-based classifiers developed in this paper achieves the best published results so far for sentence-level uncertainty prediction on the CoNLL-2010 Shared Task test data. We also show that the technique of random indexing can be successfully applied for reducing the dimensionality of the original feature space by several orders of magnitude, without sacrificing classifier performance.ConclusionsThis paper introduces a simplified approach to detecting speculation or uncertainty in text, focusing on the biomedical domain. Evaluated at the sentence-level, our SVM-based classifiers achieve the best published results so far. We also show that the feature space can be aggressively compressed using random indexing while still maintaining comparable classifier performance.

Erik Velldal | Erik Velldal

[1] Jacek M. Zurada,et al. Computational Intelligence: Imitating Life , 1994 .

[2] Thorsten Brants,et al. TnT – A Statistical Part-of-Speech Tagger , 2000, ANLP.

[3] Lyle Ungar. KDD-2006 : proceedings of the Twelfth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, August 20-23, 2006, Philadelphia, PA, USA , 2006 .

[4] Padraic Monaghan,et al. Proceedings of the 23rd annual conference of the cognitive science society , 2001 .

[5] Vladimir Vapnik,et al. The Nature of Statistical Learning , 1995 .

[6] Andreas Vlachos,et al. Detecting Speculative Language Using Syntactic Dependencies and Logistic Regression , 2010, CoNLL Shared Task.

[7] K. Bretonnel Cohen,et al. Proceedings of the BioNLP 2009 Workshop , 2009 .

[8] Magnus Sahlgren,et al. An Introduction to Random Indexing , 2005 .

[9] Sergei Nirenburg. Proceedings of the sixth conference on Applied natural language processing , 2000 .

[10] Roser Morante,et al. Learning the Scope of Hedge Cues in Biomedical Texts , 2009, BioNLP@HLT-NAACL.

[11] Beatrice Santorini,et al. Building a Large Annotated Corpus of English: The Penn Treebank , 1993, CL.