• DocumentCode
    3254399
  • Title

    Textual image compression

  • Author

    Witten, Ian H. ; Bell, Timothy C. ; Harrison, Mary-Ellen ; James, Mark L. ; Moffat, Alistair

  • Author_Institution
    Dept. of Comput. Sci., Calgary Univ., Alta., Canada
  • fYear
    1992
  • fDate
    24-27 March 1992
  • Firstpage
    42
  • Lastpage
    51
  • Abstract
    The authors describe a method for lossless compression of images that contain predominantly typed or typeset text-they call these textual images. An increasingly popular application is document archiving, where documents are scanned by a computer and stored electronically for later retrieval. Their project was motivated by such an application: Trinity College in Dublin, Ireland, are archiving their 1872 printed library catalogues onto disk, and in order to preserve the exact form of the original document, pages are being stored as scanned images rather than being converted to text. The test images are taken from this catalogue. These typeset documents have a rather old-fashioned look, and contain a wide variety of symbols from several different typefaces-the five test images used contain text in English, Flemish, Latin and Greek, and include italics and small capitals as well as roman letters. The catalogue also contains Hebrew, Syriac, and Russian text.<>
  • Keywords
    data compression; image coding; Dublin; English; Flemish; Greek; Ireland; Latin; Trinity College; document archiving; encoding; lossless compression; printed library catalogues; textural image compression; Application software; Character recognition; Computer science; Image coding; Libraries; Optical character recognition software; Power system reliability; Telephony; Testing; Typesetting;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Data Compression Conference, 1992. DCC '92.
  • Conference_Location
    Snowbird, UT, USA
  • Print_ISBN
    0-8186-2717-4
  • Type

    conf

  • DOI
    10.1109/DCC.1992.227477
  • Filename
    227477