Multilingual Document Clustering Using Wikipedia as External Knowledge

This paper presents Multilingual Document Clustering (MDC) on comparable corpora. Wikipedia has evolved to be a major structured multilingual knowledge base. It has been highly exploited in many monolingual clustering approaches and also in comparing multilingual corpora. But there is no prior work which studied the impact of Wikipedia on MDC. Here, we have studied availing Wikipedia in enhancing MDC performance. We have leveraged Wikipedia knowledge structure (such as cross-lingual links, category, outlinks, Infobox information, etc.) to enrich the document representation for clustering multilingual documents. We have implemented Bisecting k-means clustering algorithm and experiments are conducted on a standard dataset provided by FIRE for their 2010 Ad-hoc Cross-Lingual document retrieval task on Indian languages. We have considered English and Hindi datasets for our experiments. By avoiding language-specific tools, our approach provides a general framework which can be easily extendable to other languages. The system was evaluated using F-score and Purity measures and the results obtained were encouraging.

[1]  George Karypis,et al.  A Comparison of Document Clustering Techniques , 2000 .

[2]  Bruno Pouliquen,et al.  Exploiting multilingual nomenclatures and language-independent text features as an interlingua for cross-lingual text analysis applications , 2006, ArXiv.

[3]  Judith L. Klavans,et al.  Columbia Newsblaster: Multilingual News Summarization on the Web , 2004, NAACL.

[4]  Carlotta Domeniconi,et al.  Building semantic kernels for text classification using wikipedia , 2008, KDD.

[5]  Bruno Pouliquen,et al.  Cross-Lingual Document Similarity Calculation Using the Multilingual Thesaurus EUROVOC , 2002, CICLing.

[6]  Gerard Salton,et al.  A vector space model for automatic indexing , 1975, CACM.

[7]  Xiaohua Hu,et al.  Exploiting Wikipedia as external knowledge for document clustering , 2009, KDD.

[8]  Romaric Besançon,et al.  Multilingual document clusters discovery , 2004, RIAO.

[9]  G. Karypis,et al.  Criterion Functions for Document Clustering ∗ Experiments and Analysis , 2001 .

[10]  Andreas Rauber,et al.  AN ENGLISH , FRENCH , AND GERMAN VIEW OF THE RUSSIAN INFORMATION AGENCY NOVOSTI NEWS , 2010 .

[11]  Bruno Pouliquen,et al.  Multilingual and cross-lingual news topic tracking , 2004, COLING.

[12]  Lawrence J. Leftin Newsblaster Russian-English Clustering Performance Analysis , 2003 .

[13]  Soto Montalvo,et al.  Multilingual Document Clustering: An Heuristic Approach Based on Cognate Named Entities , 2006, ACL.

[14]  Prasad Pingali,et al.  Statistical Transliteration for Cross Langauge Information Retrieval using HMM alignment and CRF , 2008, IJCNLP 2008.

[15]  Hua Li,et al.  Enhancing text clustering by leveraging Wikipedia semantics , 2008, SIGIR '08.

[16]  C. A. Coelho,et al.  A STATISTICAL APPROACH FOR MULTILINGUAL DOCUMENT CLUSTERING AND TOPIC EXTRACTION FROM CLUSTERS , 2007 .

[17]  Turid Hedlund,et al.  Dictionary-Based Cross-Language Information Retrieval: Problems, Methods, and Research Findings , 2001, Information Retrieval.