• DocumentCode
    457363
  • Title

    CAPTCHA Challenge Tradeoffs: Familiarity of Strings versus Degradation of Images

  • Author

    Sui-Yu Wang ; Baird, Henry S.

  • Author_Institution
    Dept. of Comput. Sci. & Eng., Lehigh Univ., Bethlehem, PA
  • Volume
    3
  • fYear
    0
  • fDate
    0-0 0
  • Firstpage
    164
  • Lastpage
    167
  • Abstract
    It is a well documented fact that, for human readers, familiar text is more legible than unfamiliar text. Current-generation computer vision systems also are able to exploit some kinds of prior knowledge of linguistic context: for example, many OCR systems can use known lexica (word-lists, such as of commonly occurring English words) to disambiguate interpretations. It is interesting that human readers can exploit various degrees of familiarity; for example, strings of characters which, while not found in dictionaries, are similar to spelled words: e.g. "pronounceable" strings, or strings made up of frequently occurring character n-grams. In contrast to this, computer vision technologies for exploiting such poorly characterized constraints (absent an explicit, complete lexicon) are not yet well developed. This gap in ability may allow us to design stronger CAPTCHAs. We measure the familiarity of challenge strings generated by four methods (described by Bentley and Mallows) and we use the ScatterType CAPTCHA to degrade challenge images. We report the results of a human legibility trial which supports the hypothesis that more familiar strings are indeed more legible in CAPTCHAs. Our measurements may enable engineering CAPTCHAs with a more uniform distribution of difficulty by balancing image degradations against familiarity
  • Keywords
    document image processing; string matching; text analysis; OCR systems; ScatterType CAPTCHA; automated public Turing test; character n-gram; character string familiarity; computer vision system; familiar string; human legibility trial; image degradation; lexica; linguistic context; Computer science; Computer vision; Degradation; Dictionaries; Hip; Humans; Internet; Optical character recognition software; Scattering; Text analysis;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Pattern Recognition, 2006. ICPR 2006. 18th International Conference on
  • Conference_Location
    Hong Kong
  • ISSN
    1051-4651
  • Print_ISBN
    0-7695-2521-0
  • Type

    conf

  • DOI
    10.1109/ICPR.2006.355
  • Filename
    1699493