Title :
Ligature Segmentation for Urdu OCR
Author :
Lehal, Gurpreet Singh
Author_Institution :
Dept. of Comput. Sci., Punjabi Univ., Patiala, India
Abstract :
Urdu script uses superset of Arabic alphabet, but uses Nastaliq writing style. Nastaliq script is highly cursive, context sensitive and is written diagonally from top right to bottom left with stacking of characters, which makes it very hard to process for OCR. In addition, line and word segmentation are non-trivial tasks as we have frequently merging lines and vertically overlapping words and ligatures. Due to the challanges in character segmentation most of the researchers have taken the next higher unit, ligature, as recognition unit. A ligature is a connected component of one or more characters including diacritic marks and usually an Urdu word is composed of 1 to 8 ligatures. In this paper, we present a methodology for segmenting the Urdu text into ligatures. A hybrid approach, which uses top down technique for line segmentation and bottom up design for segmenting the line into ligatures, has been employed. The various challenges encountered during ligature segmentation such as horizontally overlapping and broken lines, merged ligatures and diacritic association have been discussed in detail.
Keywords :
image segmentation; natural language processing; optical character recognition; text analysis; Arabic alphabet; Nastaliq script; Nastaliq writing style; Urdu OCR; Urdu script; Urdu text segmentation; Urdu word; bottom up design; broken lines; character segmentation; diacritic association; diacritic marks; horizontally overlapping lines; ligature segmentation; line segmentation; merged ligatures; top down technique; word segmentation; Accuracy; Image color analysis; Image segmentation; Merging; Optical character recognition software; Shape; Text analysis; OCR; Urdu; ligature; segmentation;
Conference_Titel :
Document Analysis and Recognition (ICDAR), 2013 12th International Conference on
Conference_Location :
Washington, DC
DOI :
10.1109/ICDAR.2013.229