Effects of Stop Words Elimination for Arabic Information Retrieval: A Comparative Study

The effectiveness of three stop words lists for Arabic Information Retrieval---General Stoplist, Corpus- Based Stoplist, Combined Stoplist ---were investigated in this study. Three popular weighting schemes were examined: the inverse document frequency weight, probabilistic weighting, and statistical language modelling. The Idea is to combine the statistical approaches with linguistic approaches to reach an optimal performance, and compare their effect on retrieval. The LDC (Linguistic Data Consortium) Arabic Newswire data set was used with the Lemur Toolkit. The Best Match weighting scheme used in the Okapi retrieval system had the best overall performance of the three weighting algorithms used in the study, stoplists improved retrieval effectiveness especially when used with the BM25 weight. The overall performance of a general stoplist was better than the other two lists.

[1]  Gerard Salton,et al.  Term-Weighting Approaches in Automatic Text Retrieval , 1988, Inf. Process. Manag..

[2]  Lisa Ballesteros,et al.  Improving stemming for Arabic information retrieval: light stemming and co-occurrence analysis , 2002, SIGIR '02.

[3]  Ellen M. Voorhees,et al.  Overview of TREC 2001 , 2001, TREC.

[4]  S. Siegel,et al.  Nonparametric Statistics for the Behavioral Sciences , 2022, The SAGE Encyclopedia of Research Design.

[5]  Christopher J. Fox,et al.  A stop list for general text , 1989, SIGF.

[6]  Jacques Savoy,et al.  Report on the TREC 11 Experiment: Arabic, Named Page and Topic Distillation Searches , 2002, TREC.

[7]  David A. Hull Using statistical testing in the evaluation of retrieval experiments , 1993, SIGIR.

[8]  Michael McGill,et al.  Introduction to Modern Information Retrieval , 1983 .

[9]  Susan Brewer,et al.  Information storage and retrieval , 1959, ACM '59.

[10]  Leah S. Larkey,et al.  Arabic Information Retrieval at UMass in TREC-10 , 2001, TREC.

[11]  Fredric C. Gey,et al.  Translation Term Weighting and Combining Translation Resources in Cross-Language Retrieval , 2001, TREC.

[12]  Peter Schauble Multimedia Information Retrieval: Content-Based Information Retrieval from Large Text and Audio Databases , 2012 .

[13]  Abdullah M. AlShehri Optimization and effectiveness of n-grams approach for indexing and retrieval in Arabic information retrieval systems , 2002 .

[14]  Ellen M. Voorhees,et al.  The Tenth Text REtrieval Conference, TREC 2001 | NIST , 2002 .