Semantic enrichment of text representation with wikipedia for text classification

Text classification is a widely studied topic in the area of machine learning. A number of techniques have been developed to represent and classify text documents. Most of the techniques try to achieve good classification performance while taking a document only by its words (e.g. statistical analysis on word frequency and distribution patterns). One of the recent trends in text classification research is to incorporate more semantic interpretation in text classification, especially by using Wikipedia. This paper introduces a technique for incorporating the vast amount of human knowledge accumulated in Wikipedia into text representation and classification. The aim is to improve classification performance by transforming general terms into a set of related concepts grouped around semantic themes. In order to achieve this goal, this paper proposes a unique method for breaking the enormous amount of extracted Wikipedia knowledge (concepts) into smaller pieces (subsets of concepts). The subsets of concepts are separately used to represent the same set of documents in a number of different ways, from which an ensemble of classifiers is built. Experimental results show that an ensemble of classifiers individually trained on a different representation of the document set performs better with increased accuracy and stability than that of a classifier trained only on the original document set.

[1]  Christos Makris,et al.  k-Attractors: A Clustering Algorithm for Software Measurement Data Analysis , 2007 .

[2]  Takahiro Hara,et al.  Association thesaurus construction methods based on link co-occurrence analysis for wikipedia , 2008, CIKM '08.

[3]  Louis Vuurpijl,et al.  An overview and comparison of voting methods for pattern recognition , 2002, Proceedings Eighth International Workshop on Frontiers in Handwriting Recognition.

[4]  Evgeniy Gabrilovich,et al.  Wikipedia-based Semantic Interpretation for Natural Language Processing , 2014, J. Artif. Intell. Res..

[5]  Brian Vickery Information Representation and Retrieval in the Digital Age , 2004, Program.

[6]  Steven J. Simske,et al.  Performance analysis of pattern classifier combination by plurality voting , 2003, Pattern Recognit. Lett..

[7]  R. Polikar,et al.  Ensemble based systems in decision making , 2006, IEEE Circuits and Systems Magazine.

[8]  Peter W. Foltz,et al.  An introduction to latent semantic analysis , 1998 .

[9]  Patrick F. Reidy An Introduction to Latent Semantic Analysis , 2009 .

[10]  Gerhard Weikum,et al.  Learning Word-to-Concept Mappings for Automatic Text Classification , 2005, ICML 2005.

[11]  Angela Heath,et al.  Information representation and retrieval in the digital age , 2005, J. Assoc. Inf. Sci. Technol..

[12]  Xiaohua Hu,et al.  Dragon Toolkit: Incorporating Auto-Learned Semantic Knowledge into Large-Scale Text Retrieval and Mining , 2007, 19th IEEE International Conference on Tools with Artificial Intelligence(ICTAI 2007).

[13]  Gerard Salton,et al.  A vector space model for automatic indexing , 1975, CACM.