论文信息 - Document clustering using sequential pattern (SP): Maximal frequent sequences (MFS) as SP representation

Document clustering using sequential pattern (SP): Maximal frequent sequences (MFS) as SP representation

This research proposes an idea to apply Feature Based Clustering (FBC) in document clustering. A huge number of existing documents will be easier to be used if they are clustered into several topics. FBC uses K-Means algorithm to cluster sequential data of features. Features of text document can be presented as sequence of word. In order to be processed as sequential data, features must be extracted from collection of unstructured text documents. Therefore, we need preprocessing tasks to deliver appropriate form of document features. There are two types of sequential pattern using simple form: Frequent Word Sequence (FWS) and Maximal Frequent Sequence (MFS). Both types are appropriate for text data. The difference is in applying the maximum principle in MFS. Therefore, MFS amount from a text document would be less than the amount of its FWS. In this research, we choose maximal frequent sequences (MFS) as feature representation. We proposes framework to conduct FBC using MFS as features. The framework is tested to cluster dataset that is subset of the Twenty News Group Text Data. The result shows that the accuracy of clustering result is affected by the parameter's value, dataset, and the number of target cluster.

G. A. Putri Saptawati | Yani Widyani | Dini Rahmawati

[1] Anne Laurent,et al. Sequential patterns for text categorization , 2006, Intell. Data Anal..

[2] Helena Ahonen-Myka. Mining all maximal frequent word sequences in a set of sentences , 2005, CIKM '05.

[3] Helena Ahonen. Knowledge Discovery in Documents by Extracting Frequent Word Sequences , 1999, Libr. Trends.

[4] Wei-Ying Ma,et al. Multi-type Features Based Web Document Clustering , 2004, WISE.

[5] Antoine Doucet,et al. Advanced document description, a sequential approach , 2006, SIGF.

[6] Elisa Margareth Sibarani. STUDI DAN IMPLEMENTASI TEKNIK SEMI-SUPERVISED CLUSTERING BERBASIS EXPECTATION MAXIMIZAATION (EM) UNTUK DOCUMENT CLUSTERING , 2009 .

[7] Mohamed S. Kamel,et al. Efficient phrase-based document indexing for Web document clustering , 2004, IEEE Transactions on Knowledge and Data Engineering.

[8] Jiawei Han,et al. Data Mining: Concepts and Techniques , 2000 .

[9] Valerie Guralnik,et al. A scalable algorithm for clustering sequential data , 2001, Proceedings 2001 IEEE International Conference on Data Mining.