Syntactic Segmentation and Labeling of Digitized Pages from Technical Journals

A method for extracting alternating horizontal and vertical projection profiles are from nested sub-blocks of scanned page images of technical documents is discussed. The thresholded profile strings are parsed using the compiler utilities Lex and Yacc. The significant document components are demarcated and identified by the recursive application of block grammars. Backtracking for error recovery and branch and bound for maximum-area labeling are implemented with Unix Shell programs. Results of the segmentation and labeling process are stored in a labeled x-y tree. It is shown that families of technical documents that share the same layout conventions can be readily analyzed. Results from experiments in which more than 20 types of document entities were identified in sample pages from two journals are presented. >

[1]  Murray Hill,et al.  Yacc: Yet Another Compiler-Compiler , 1978 .

[2]  Spyros S. Magliveras,et al.  The Number of Tilings of a Block with Blocks , 1988, Eur. J. Comb..

[3]  George Nagy,et al.  Towards a Structured-Document-Image Utility , 1992 .

[4]  George Nagy,et al.  DOCUMENT ANALYSIS WITH AN EXPERT SYSTEM , 1986 .

[5]  Kazuhiko Yamamoto,et al.  Structured Document Image Analysis , 1992, Springer Berlin Heidelberg.

[6]  Mahesh Viswanathan,et al.  A prototype document image analysis system for technical journals , 1992, Computer.

[7]  Andreas Dengel,et al.  ANASTASIL: A Hybrid Knowledge-Based System for Document Layout Analysis , 1989, IJCAI.

[8]  George Nagy,et al.  HIERARCHICAL REPRESENTATION OF OPTICALLY SCANNED DOCUMENTS , 1984 .

[9]  Eiichi Tanaka,et al.  Theoretical aspects of syntactic pattern recognition , 1995, Pattern Recognit..

[10]  George Nagy,et al.  A syntactic approach to document segmentation and labeling , 1990 .

[11]  Mahesh Viswanathan Analysis of Scanned Documents — a Syntactic Approach , 1992 .

[12]  Henry S. Baird,et al.  The skew angle of printed documents , 1995 .

[13]  Mahesh Viswanathan,et al.  A SYNTACTIC APPROACH TO DOCUMENT SEGMENTATION , 1990 .

[14]  Robert S. Ledley,et al.  Special issue on Optical Character Recognition , 1970, Pattern Recognit..

[15]  George Nagy,et al.  An Interactive System for Reading Unformatted Printed Text , 1971, IEEE Transactions on Computers.

[16]  George Nagy,et al.  Characteristics of digitized images of technical articles , 1992, Electronic Imaging.

[17]  S. V. Rice A report on the accuracy of OCR devices , 1992 .

[18]  Osamu Hori,et al.  A robust recognition system for a drawing superimposed on a map , 1992, Computer.

[19]  Mahesh Viswanathan,et al.  Matrix Quantization of Homomorphically Processed Images , 1986, MILCOM 1986 - IEEE Military Communications Conference: Communications-Computers: Teamed for the 90's.

[20]  E. Schmidt,et al.  Lex—a lexical analyzer generator , 1990 .