مرکز منطقه ای اطلاع رساني علوم و فناوري

DocumentCode :

3489256

Title :

Ligature Segmentation for Urdu OCR

Author :

Lehal, Gurpreet Singh

Author_Institution :

Dept. of Comput. Sci., Punjabi Univ., Patiala, India

fYear :

2013

fDate :

25-28 Aug. 2013

Firstpage :

1130

Lastpage :

1134

Abstract :

Urdu script uses superset of Arabic alphabet, but uses Nastaliq writing style. Nastaliq script is highly cursive, context sensitive and is written diagonally from top right to bottom left with stacking of characters, which makes it very hard to process for OCR. In addition, line and word segmentation are non-trivial tasks as we have frequently merging lines and vertically overlapping words and ligatures. Due to the challanges in character segmentation most of the researchers have taken the next higher unit, ligature, as recognition unit. A ligature is a connected component of one or more characters including diacritic marks and usually an Urdu word is composed of 1 to 8 ligatures. In this paper, we present a methodology for segmenting the Urdu text into ligatures. A hybrid approach, which uses top down technique for line segmentation and bottom up design for segmenting the line into ligatures, has been employed. The various challenges encountered during ligature segmentation such as horizontally overlapping and broken lines, merged ligatures and diacritic association have been discussed in detail.

Keywords :

image segmentation; natural language processing; optical character recognition; text analysis; Arabic alphabet; Nastaliq script; Nastaliq writing style; Urdu OCR; Urdu script; Urdu text segmentation; Urdu word; bottom up design; broken lines; character segmentation; diacritic association; diacritic marks; horizontally overlapping lines; ligature segmentation; line segmentation; merged ligatures; top down technique; word segmentation; Accuracy; Image color analysis; Image segmentation; Merging; Optical character recognition software; Shape; Text analysis; OCR; Urdu; ligature; segmentation;

fLanguage :

English

Publisher :

ieee

Conference_Titel :

Document Analysis and Recognition (ICDAR), 2013 12th International Conference on

Conference_Location :

Washington, DC

ISSN :

1520-5363

Type :

conf

DOI :

10.1109/ICDAR.2013.229

Filename :

6628790

Link To Document :

https://search.ricest.ac.ir/dl/search/defaultta.aspx?DTC=49&DC=3489256