A Novel Modified Apriori Approach for Web Document Clustering

The Traditional apriori algorithm can be used for clustering the web documents based on the association technique of data mining. But this algorithm has several limitations due to repeated database scans and its weak association rule analysis. In modern world of large databases, efficiency of traditional apriori algorithm would reduce manifolds. In this paper, we proposed a new modified apriori approach by cutting down the repeated database scans and improving association analysis of traditional apriori algorithm to cluster the web documents. Further we improve those clusters by applying Fuzzy C-Means (FCM), K-Means and Vector Space Model (VSM) techniques separately. We use Classic3 and Classic4 datasets of Cornell University having more than 10,000 documents and run both traditional apriori and our modified apriori approach on it. Experimental results show that our approach outperforms the traditional apriori algorithm in terms of database scan and improvement on association of analysis.

[1]  Heikki Mannila,et al.  Fast Discovery of Association Rules in Large Databases , 1996, Knowledge Discovery and Data Mining.

[2]  Vipin Kumar,et al.  Scalable parallel data mining for association rules , 1997, SIGMOD '97.

[3]  Ferenc Bodon,et al.  A trie-based APRIORI implementation for mining frequent item sequences , 2005 .

[4]  Heikki Mannila,et al.  Fast Discovery of Association Rules , 1996, Advances in Knowledge Discovery and Data Mining.

[5]  Yanping Li,et al.  Web Clustering Using a Two-Layer Approach , 2011, WISM.

[6]  Lin Shi,et al.  Mining Association Rules Based on Apriori Algorithm and Application , 2009, 2009 International Forum on Computer Science-Technology and Applications.

[7]  Xiaojun Cao An Algorithm of Mining Association Rules Based on Granular Computing , 2012 .

[8]  Ying Li,et al.  DM Data Mining Based on Improved Apriori Algorithm , 2013, ICICA.

[9]  Jean-Raymond Abrial,et al.  On B , 1998, B.

[10]  Xindong Wu,et al.  Clustering web documents using hierarchical representation with multi-granularity , 2012, World Wide Web.

[11]  Rajendra Kumar Roul,et al.  Web Document Clustering and Ranking using Tf-Idf based Apriori Approach , 2014, ArXiv.

[12]  Salvatore Orlando,et al.  Enhancing the Apriori Algorithm for Frequent Set Counting , 2001, DaWaK.

[13]  Jin Shang,et al.  A Novel Apriori Algorithm Based on Cross Linker , 2013 .

[14]  Daniel T. Larose,et al.  Discovering Knowledge in Data: An Introduction to Data Mining , 2005 .

[15]  Jiri Kupka,et al.  Implementation of Background Knowledge and Properties Induced by Fuzzy Confirmation Measures in Apriori Algorithm , 2012, CISIS/ICEUTE/SOCO Special Sessions.

[16]  Jaishree Singh,et al.  Improving Efficiency of Apriori Algorithm Using Transaction Reduction , 2013 .

[17]  Byung-Won On,et al.  An effective web document clustering algorithm based on bisection and merge , 2011, Artificial Intelligence Review.