论文信息 - Host-IP Clustering Technique for Deep Web Characterization

Host-IP Clustering Technique for Deep Web Characterization

A huge portion of today’s Web consists of web pages filled with information from myriads of online databases. This part of the Web, known as the deep Web, is to date relatively unexplored and even major characteristics such as number of searchable databases on the Web is somewhat disputable. In this paper, we are aimed at more accurate estimation of main parameters of the deep Web by sampling one national web domain. We propose the Host-IP clustering sampling technique that addresses drawbacks of existing approaches to characterize the deep Web and report our findings based on the survey of Russian Web conducted in September 2006. Obtained estimates together with a proposed sampling method could be useful for further studies to handle data in the deep Web.

Tapio Salakoski | Denis Shestakov

[1] Denis Shestakov,et al. Deep Web: Databases on the Web , 2009 .

[2] Edward T. O'Neill,et al. A Methodology for Sampling the World Wide Web , 2001 .

[3] B. Huberman,et al. The Deep Web : Surfacing Hidden Value , 2000 .

[4] Denis Shestakov. On Building a Search Interface Discovery System , 2009, RED.

[5] Ricardo A. Baeza-Yates,et al. Characterization of national Web domains , 2007, TOIT.

[6] Tapio Salakoski,et al. On Estimating the Scale of National Deep Web , 2007, DEXA.

[7] Mitesh Patel,et al. Accessing the deep web , 2007, CACM.

[8] Andrei Z. Broder,et al. A Technique for Measuring the Relative Size and Overlap of Public Web Search Engines , 1998, Comput. Networks.