论文信息 - Collocation Extraction Using Web Statistics

Collocation Extraction Using Web Statistics

This paper mines collocations from two different web usage corpora, NTU proxy log and TTS search log. The precisions for NTU and TTS test data are 61.76% and 57.50%, respectively, by human judgment for 2% sampling of extracted collocations. For automatic evaluation, we submit extracted collocation to Google search engine, and the resulting page counts are used to compute the mutual information of the collocation. Experimental results show that total 43.27% and 42.65% of collocations mined from NTU and TTS corpora passed the examination of MIs.

Hsin-Hsi Chen | Chih-Long Lin | Yi-Cheng Yu

[1] Yen-Jen Oyang,et al. A Contextual Term Suggestion Mechanism for Interactive Web Search , 2001, Web Intelligence.

[2] Mark Hansen,et al. Using navigation data to improve IR functions in the context of web search , 2001, CIKM '01.

[3] Jaideep Srivastava,et al. Web usage mining: discovery and applications of usage patterns from Web data , 2000, SKDD.

[4] Wei-Ying Ma,et al. Probabilistic query expansion using query logs , 2002, WWW '02.