论文信息 - Experiments with a component theory of probabilistic information retrieval based on single terms as document components

Experiments with a component theory of probabilistic information retrieval based on single terms as document components

A component theory of information retrieval using single content terms as component for queries and documents was reviewed and experimented with. The theory has the advantages of being able to (1) bootstrap itself, that is, define initial term weights naturally based on the fact that items are self relevent; (2) make use of within-item term frequencies; (3) account for query-focused and document-focused indexing and retrieval strategies cooperatively; and (4) allow for component-specific feedback if such information is available. Retrieval results with four collections support the effectiveness of all the first three aspects, except for predictive retrieval. At the initial indexing stage, the retrieval theory performed much more consistantly across collections than croft's model and provided results comparable to Salton's tf*idf approach. An inverse collection term frequency (ICTF) formula was also tested that performed much better than the inverse document frequency (IDF). With full feedback retrospective retrieval, the component theory performed substantially better than Croft's, because of the highly specific nature of document-focused feedback. Repetitive retireval results with partial relevance feedback mirrored those for the retrospective. However, for the important case of predictive retrieval using residual ranking, results were not unequivocal.

Kui-Lam Kwok | K. Kwok

[1] Stephen P. Harter,et al. A probabilistic approach to automatic keyword indexing. Part II. An algorithm for probabilistic indexing , 1975, J. Am. Soc. Inf. Sci..

[2] Edward A. Fox,et al. Development of the coder system: A testbed for artificial intelligence methods in information retrieval , 1987, Inf. Process. Manag..

[3] Alan F. Smeaton,et al. The Retrieval Effects of Query Expansion on a Feedback Document Retrieval System , 1983, Comput. J..

[4] Clement T. Yu,et al. On the Construction of Feedback Queries , 1982, JACM.

[5] Clement T. Yu,et al. A framework for effective retrieval , 1989, ACM Trans. Database Syst..

[6] Kui-Lam Kwok. An interpretation of index term weighting schemes based on document components , 1986, SIGIR '86.

[7] Clement T. Yu,et al. The measurement of term importance in automatic indexing , 1981, J. Am. Soc. Inf. Sci..

[8] Stephen E. Robertson,et al. Probabilistic models of indexing and searching , 1980, SIGIR '80.

[9] W. Bruce Croft,et al. I3R: A new approach to the design of document retrieval systems , 1987, J. Am. Soc. Inf. Sci..

[10] D. A. Kemp. Relevance, pertinence and information system development , 1974, Inf. Storage Retr..

[11] SaltonGerard,et al. Term-weighting approaches in automatic text retrieval , 1988 .