• DocumentCode
    13597
  • Title

    Language-Independent Text-Line Extraction Algorithm for Handwritten Documents

  • Author

    Jewoong Ryu ; Hyung Il Koo ; Nam Ik Cho

  • Author_Institution
    Dept. of Electr. & Comput. Eng., Seoul Nat. Univ., Seoul, South Korea
  • Volume
    21
  • Issue
    9
  • fYear
    2014
  • fDate
    Sept. 2014
  • Firstpage
    1115
  • Lastpage
    1119
  • Abstract
    Text-line extraction in handwritten documents is an important step for document image understanding, and a number of algorithms have been proposed to address this problem. However, most of them exploit features of specific languages and work only for a given language. In order to overcome this limitation, we develop a language-independent text-line extraction algorithm. Our method is based on connected components (CCs), however, unlike conventional methods, we analyze strokes and partition under-segmented CCs into normalized ones. Due to this normalization, the proposed method is able to estimate the states of CCs for a range of different languages and writing styles. From the estimated states, we build a cost function whose minimization yields text-lines. Experimental results show that the proposed method yields the state-of-the-art performance on Latin-based and Chinese script databases. Further, we submitted the proposed algorithm to the ICDAR 2013 handwriting segmentation competition and our method showed the best text-line extraction performance among 10 participant methods.
  • Keywords
    document image processing; handwritten character recognition; image segmentation; text analysis; ICDAR 2013 handwriting segmentation competition; Latin-based and Chinese script databases; components; cost function; document image understanding; handwritten documents; language-independent text-line extraction algorithm; partition under-segmented CCs; writing styles; yield text-line minimization; Cost function; Data mining; Databases; Feature extraction; Minimization; Partitioning algorithms; Signal processing algorithms; Connected component based algorithm; handwritten documents; language-independent algorithm; text-line extraction; text-line segmentation;
  • fLanguage
    English
  • Journal_Title
    Signal Processing Letters, IEEE
  • Publisher
    ieee
  • ISSN
    1070-9908
  • Type

    jour

  • DOI
    10.1109/LSP.2014.2325940
  • Filename
    6819023