• Title of article

    Text Extraction in Complex Color Document Images for Enhanced Readability

  • Author/Authors

    P. Nagabhushan، نويسنده , , S. Nirmala، نويسنده ,

  • Issue Information
    روزنامه با شماره پیاپی سال 2010
  • Pages
    14
  • From page
    120
  • To page
    133
  • Abstract
    Often we encounter documents with text printed on complex color background. Readability of textual con-tents in such documents is very poor due to complexity of the background and mix up of color(s) of fore-ground text with colors of background. Automatic segmentation of foreground text in such document images is very much essential for smooth reading of the document contents either by human or by machine. In this paper we propose a novel approach to extract the foreground text in color document images having complex background. The proposed approach is a hybrid approach which combines connected component and texture feature analysis of potential text regions. The proposed approach utilizes Canny edge detector to detect all possible text edge pixels. Connected component analysis is performed on these edge pixels to identify can-didate text regions. Because of background complexity it is also possible that a non-text region may be iden-tified as a text region. This problem is overcome by analyzing the texture features of potential text region corresponding to each connected component. An unsupervised local thresholding is devised to perform fore-ground segmentation in detected text regions. Finally the text regions which are noisy are identified and re-processed to further enhance the quality of retrieved foreground. The proposed approach can handle docu-ment images with varying background of multiple colors and texture; and foreground text in any color, font, size and orientation. Experimental results show that the proposed algorithm detects on an average 97.12% of text regions in the source document. Readability of the extracted foreground text is illustrated through Opti-cal character recognition (OCR) in case the text is in English. The proposed approach is compared with some existing methods of foreground separation in document images. Experimental results show that our approach performs better.
  • Keywords
    OCR , Complex background , Connected Component Analysis , Segmentation of Text , texture analysis , Unsupervised Thresholding , Color Document Image
  • Journal title
    Intelligent Information Management
  • Serial Year
    2010
  • Journal title
    Intelligent Information Management
  • Record number

    664381