论文信息 - The importance of anchor text for ad hoc search revisited

The importance of anchor text for ad hoc search revisited

It is generally believed that propagated anchor text is very important for effective Web search as offered by the commercial search engines. "Google Bombs" are a notable illustration of this. However, many years of TREC Web retrieval research failed to establish the effectiveness of link evidence for ad hoc retrieval on Web collections. The ultimate resolution to this dilemma was that typical Web search is very different from the traditional ad hoc methodology. So far, however, no one has established why link information, like incoming link degree or anchor text, does not help ad hoc retrieval effectiveness. Several possible explanations were given, including the collections being too small for anchors to be effective, and the density of the link graph being too low. The new TREC 2009 Web Track collection is substantially larger than previous collections and has a dense link graph. Our main finding is that propagated anchor text outperforms full-text retrieval in terms of early precision, and in combination with it, gives an improvement in overall precision. We then analyse the impact of link density and collection size by down-sampling the number of links and the number of pages respectively. Other findings are that, contrary to expectations, (inter-server) link density has little impact on effectiveness, while the size of the collection has a substantial impact on the quantity, quality and effectiveness of anchor text. We also compare the diversity of the search results of anchor text and full-text approaches, which show that anchor text performs significantly better than full-text search and confirm our findings for the ad hoc search task.

Jaap Kamps | Marijn Koolen

[1] Peter Bailey,et al. Is it fair to evaluate Web systems using TREC ad hoc methods , 1999, SIGIR 1999.

[2] Kevin S. McCurley,et al. Analysis of anchor text for web search , 2003, SIGIR.

[3] David Hawking,et al. How Valuable is External Link Evidence When Searching Enterprise Webs? , 2004, ADC.

[4] Stephen E. Robertson,et al. Effective site finding using link anchor information , 2001, SIGIR '01.

[5] Nick Craswell,et al. The impact of crawl policy on web search effectiveness , 2009, SIGIR.

[6] Amit Singhal,et al. A case study in web search using TREC algorithms , 2001, WWW '01.

[7] David Hawking,et al. Overview of the TREC-9 Web Track , 2000, TREC.

[8] Stephen E. Robertson,et al. On Collection Size and Retrieval Effectiveness , 2004, Information Retrieval.

[9] James P. Callan,et al. Combining document representations for known-item search , 2003, SIGIR.

[10] Peter Bailey,et al. Overview of the TREC-8 Web Track , 2000, TREC.

[11] Serge Abiteboul,et al. Adaptive on-line page importance computation , 2003, WWW '03.