论文信息 - An Extension of the VSM Documents Representation

An Extension of the VSM Documents Representation

In this paper we will present a new approach regarding the documents representation in order to be used in classification and/or clustering algorithms. In our new representation we will start from the classical "bag-of-words" representation but we will augment each word with its correspondent part-of-speech. Thus we will introduce a new concept called hyper-vectors where each document is represented in a hyper-space where each dimension is a different part-of-speech component. For each dimension the document is represented using the Vector Space Model (VSM). In this work we will use only five different parts of speech: noun, verb, adverb, adjective and others. In the hyper-space each dimension has a different weight. To compute the similarity between two documents we have developed a new hyper-cosine formula. Some interesting classification experiments are presented as validation cases.

[1] Ruslan Mitkov,et al. The Oxford handbook of computational linguistics , 2003 .

[2] Ines Gloeckner. Mining The Web Discovering Knowledge From Hypertext Data , 2016 .

[3] A. David,et al. Part-of-speech labeling for Reuters database , 2015, 2015 19th International Conference on System Theory, Control and Computing (ICSTCC).

[4] A. David,et al. Part of speech tagging with Naïve Bayes methods , 2014, 2014 18th International Conference on System Theory, Control and Computing (ICSTCC).

[5] R. Suganya,et al. Data Mining Concepts and Techniques , 2010 .