This paper proposes a general methodology for automatic layout segmentation of documents. We first use colour histograms for extracting dominant colours of an image. This information is then used to hierarchically segment documents into regions of interest represented as polygons. If a region of interest is a picture the algorithm intelligently refrains from segmenting it further, while coloured regions that contain text are subsegmented. The method has been tested on 50 real life documents, such as office letters, brochures, and technical papers, scanned at 100/spl times/100 dpi resolution. Regions are detected with about 68% reliability. A critical analysis of the results is presented.
展开▼