A Note on Term Weighting and Text Matching
暂无分享,去创建一个
In information retrieval, it is not uncommon to be faced with large collections of unrestricted natural-language text. In such circumstances, the text analysis and retrieval operations must be based mainly on a study of the text collections actually under construction. Two main operations are of interest: a text analysis operation designed to assign content identifiers to the stored texts, and a text comparison system designed to identify texts covering particular subject areas. In the present note, some details are given concerning the usefulness of term weighting systems for the content analysis of natural-language texts, and of text matching strategies designed to identify relevant text items in answer to available search requests. A sample collection of electronic mail messages is used for experimental purposes.