DocumentCode
3254399
Title
Textual image compression
Author
Witten, Ian H. ; Bell, Timothy C. ; Harrison, Mary-Ellen ; James, Mark L. ; Moffat, Alistair
Author_Institution
Dept. of Comput. Sci., Calgary Univ., Alta., Canada
fYear
1992
fDate
24-27 March 1992
Firstpage
42
Lastpage
51
Abstract
The authors describe a method for lossless compression of images that contain predominantly typed or typeset text-they call these textual images. An increasingly popular application is document archiving, where documents are scanned by a computer and stored electronically for later retrieval. Their project was motivated by such an application: Trinity College in Dublin, Ireland, are archiving their 1872 printed library catalogues onto disk, and in order to preserve the exact form of the original document, pages are being stored as scanned images rather than being converted to text. The test images are taken from this catalogue. These typeset documents have a rather old-fashioned look, and contain a wide variety of symbols from several different typefaces-the five test images used contain text in English, Flemish, Latin and Greek, and include italics and small capitals as well as roman letters. The catalogue also contains Hebrew, Syriac, and Russian text.<>
Keywords
data compression; image coding; Dublin; English; Flemish; Greek; Ireland; Latin; Trinity College; document archiving; encoding; lossless compression; printed library catalogues; textural image compression; Application software; Character recognition; Computer science; Image coding; Libraries; Optical character recognition software; Power system reliability; Telephony; Testing; Typesetting;
fLanguage
English
Publisher
ieee
Conference_Titel
Data Compression Conference, 1992. DCC '92.
Conference_Location
Snowbird, UT, USA
Print_ISBN
0-8186-2717-4
Type
conf
DOI
10.1109/DCC.1992.227477
Filename
227477
Link To Document