Geometric Algorithms and Experiments for Automated Document Structuring

We present and analyze algorithms for the automated segmentation and classiication of layout structures in electronic documents. The key idea is to use the patterns in the distribution of white space in a document to recognize and interpret its components. The segmentation algorithm divides the document into a hierarchy of logical elements; the clas-siication algorithms classify these divisions as base-text, tables, indented lists, polygonal drawings, and graphs. We present experimental data and discuss an information access application. Our methodology allows the automatic markup of documents (for instance in the sgml format) and the creation of multi-level indices and browsing tools for electronic libraries.

[1]  Masaaki Mizuno,et al.  Document Recognition System with Layout Structure Generator , 1990, MVA.

[2]  田中 洋子 "Daniel Deronda"論 , 1970 .

[3]  Charles F. Goldfarb,et al.  SGML handbook , 1990 .

[4]  Marti A. Hearst Contextualizing Retrieval of Full-Length Documents , 1994 .

[5]  Gerard Salton,et al.  Automatic Text Processing: The Transformation, Analysis, and Retrieval of Information by Computer , 1989 .

[6]  Devika Subramanian,et al.  Multi-media RISC informatics: retrieving information with simple structural components , 1993, CIKM '93.

[7]  K. S. Baird,et al.  Anatomy of a versatile page reader , 1992, Proc. IEEE.

[8]  James Allan,et al.  Information Agents for Building Hyperlinks , 1993 .

[9]  Zhigang Fan,et al.  Tabular document recognition , 1994, Electronic Imaging.

[10]  Theo Huibers Detecting the erosion of hierarchic information structures , 1993 .

[11]  Jeffrey D. Smith,et al.  Design and Analysis of Algorithms , 2009, Lecture Notes in Computer Science.

[12]  David R. Karger,et al.  Scatter/Gather: a cluster-based approach to browsing large document collections , 1992, SIGIR '92.

[13]  Leslie Lamport,et al.  Latex : A Document Preparation System , 1985 .

[14]  Haruo Asada,et al.  Major components of a complete text reading system , 1992 .

[15]  Anil K. Jain,et al.  Address block location on envelopes using Gabor filters , 1992, Pattern Recognit..

[16]  Mahesh Viswanathan,et al.  A prototype document image analysis system for technical journals , 1992, Computer.