The availability of online text documents exposes readers to a vast amount of potentially valuable knowledge buried therein. The sheer scale of material has created the pressing need for automated methods of discovering relevant information without having to read it all. Hence the growing interest in recent years in Text Mining.A common approach to Text Mining is Information Extraction (IE), extracting specific types (or templates) of information from a document collection. Although many works on IE have been published, researchers have not paid much attention to evaluate the contribution of syntactic and semantic analysis using Natural Language Processing (NLP) techniques to the quality of IE results.In this work we try to quantify the contribution of NLP techniques, by comparing three strategies for IE: naive co-occurrence, ordered co-occurrence, and the structure-driven method - a rule-based strategy that relies on syntactic analysis followed by the extraction of suitable semantic templates. We use the three strategies for the extraction of two templates from financial news stories. We show that the structure-driven strategy provides significantly better precision results than the two other strategies (80-90% for the structure-driven compared with about only 60% for the co-occurrence and ordered co-occurrence). These results indicate that a syntactical and semantic analysis is necessary if one wishes to obtain high accuracy.
[1]
Dekang Lin,et al.
University of Manitoba: Description of the PIE System Used for MUC-6
,
1995,
MUC.
[2]
Jonathan Aseltine.
WAVE: An Incremental Algorithm for Information Extraction
,
1999
.
[3]
Douglas E. Appelt,et al.
FASTUS: A Finite-state Processor for Information Extraction from Real-world Text
,
1993,
IJCAI.
[4]
Stephen Glenn Soderland,et al.
Learning text analysis rules for domain-specific natural language processing
,
1996
.
[5]
David Fisher,et al.
Automatically Learned vs. Hand-crafted Text Analysis Rules
,
1997
.
[6]
Ronen Feldman,et al.
A framework for specifying explicit bias for revision of approximate information extraction rules
,
2000,
KDD '00.