SKIF-P: a point-based indexing and ranking of web documents for spatial-keyword search

There is a significant commercial and research interest in location-based web search engines. Given a number of search keywords and one or more locations (geographical points) that a user is interested in, a location-based web search retrieves and ranks the most textually and spatially relevant web pages. In this type of search, both the spatial and textual information should be indexed. Currently, no efficient index structure exists that can handle both the spatial and textual aspects of data simultaneously and accurately. Existing approaches either index space and text separately or use inefficient hybrid index structures with poor performance and inaccurate results. Moreover, most of these approaches cannot accurately rank web-pages based on a combination of space and text and are not easy to integrate into existing search engines. In this paper, we propose a new index structure called Spatial-Keyword Inverted File for Points to handle point-based indexing of web documents in an integrated/efficient manner. To seamlessly find and rank relevant documents, we develop a new distance measure called spatial tf-idf. We propose four variants of spatial-keyword relevance scores and two algorithms to perform top-k searches. As verified by experiments, our proposed techniques outperform existing index structures in terms of search performance and accuracy.

[1]  Ming Du,et al.  gR*-tree: An Index for Querying Approximate Keywords in Geographic Information System , 2009, 2009 International Conference on Information Engineering and Computer Science.

[2]  Mark Sanderson,et al.  Spatio-textual Indexing for Geographical Search on the Web , 2005, SSTD.

[3]  Chen Li,et al.  Processing Spatial-Keyword (SK) Queries in Geographic Information Retrieval (GIR) Systems , 2007, 19th International Conference on Scientific and Statistical Database Management (SSDBM 2007).

[4]  Chen Li,et al.  Supporting location-based approximate-keyword queries , 2010, GIS '10.

[5]  Esther M. Arkin,et al.  Approximations for minimum and min-max vehicle routing problems , 2006, J. Algorithms.

[6]  Naphtali Rishe,et al.  Keyword Search on Spatial Databases , 2008, 2008 IEEE 24th International Conference on Data Engineering.

[7]  Christian S. Jensen,et al.  Retrieving top-k prestige-based relevant spatial web objects , 2010, Proc. VLDB Endow..

[8]  Kevin S. McCurley,et al.  Geospatial mapping and navigation of the web , 2001, WWW '01.

[9]  Hinrich Schütze,et al.  Introduction to information retrieval , 2008 .

[10]  Torsten Suel,et al.  Three-Level Caching for Efficient Query Processing in Large Web Search Engines , 2005, WWW '05.

[11]  W. Tobler A Computer Movie Simulating Urban Growth in the Detroit Region , 1970 .

[12]  Torsten Suel,et al.  Efficient query processing in geographic web search engines , 2006, SIGMOD Conference.

[13]  Divesh Srivastava,et al.  Forward Decay: A Practical Time Decay Model for Streaming Systems , 2009, 2009 IEEE 25th International Conference on Data Engineering.

[14]  Taher H. Haveliwala Topic-sensitive PageRank , 2002, IEEE Trans. Knowl. Data Eng..

[15]  Xing Xie,et al.  Hybrid index structures for location-based web search , 2005, CIKM '05.

[16]  Hyun Chul Lee,et al.  Geographically focused collaborative crawling , 2006, WWW '06.

[17]  Christian S. Jensen,et al.  Efficient Retrieval of the Top-k Most Relevant Spatial Web Objects , 2009, Proc. VLDB Endow..

[18]  Torsten Suel,et al.  Three-level caching for efficient query processing in large Web search engines , 2005, WWW.

[19]  Peter Willett,et al.  Readings in information retrieval , 1997 .

[20]  Chen Li,et al.  Hybrid Indexing and Seamless Ranking of Spatial and Textual Features of Web Documents , 2010, DEXA.

[21]  SaltonGerard,et al.  Term-weighting approaches in automatic text retrieval , 1988 .

[22]  Alistair Moffat,et al.  Adding compression to a full‐text retrieval system , 1995, Softw. Pract. Exp..

[23]  Anthony K. H. Tung,et al.  Keyword Search in Spatial Databases: Towards Searching by Document , 2009, 2009 IEEE 25th International Conference on Data Engineering.

[24]  Luis Gravano,et al.  Computing Geographical Scopes of Web Resources , 2000, VLDB.

[25]  JUSTIN ZOBEL,et al.  Inverted files for text search engines , 2006, CSUR.

[26]  Ron Sivan,et al.  Web-a-where: geotagging web content , 2004, SIGIR '04.

[27]  Edith Cohen,et al.  Maintaining time-decaying stream aggregates , 2006, J. Algorithms.

[28]  Byeong-Soo Jeong,et al.  Inverted File Partitioning Schemes in Multiple Disk Systems , 1995, IEEE Trans. Parallel Distributed Syst..