Using Word Sequences for Text Summarization

Traditional approaches for extractive summarization score/classify sentences based on features such as position in the text, word frequency and cue phrases These features tend to produce satisfactory summaries, but have the inconvenience of being domain dependent In this paper, we propose to tackle this problem representing the sentences by word sequences (n-grams), a widely used representation in text categorization The experiments demonstrated that this simple representation not only diminishes the domain and language dependency but also enhances the summarization performance.