• DocumentCode
    153412
  • Title

    AreCAPTCHA: Outsourcing Arabic Text Digitization to Native Speakers

  • Author

    Bakry, Menna ; Khamis, Mohamed ; Abdennadher, Slim

  • Author_Institution
    Comput. Sci. & Eng. Dept., German Univ. in Cairo, Cairo, Egypt
  • fYear
    2014
  • fDate
    7-10 April 2014
  • Firstpage
    304
  • Lastpage
    308
  • Abstract
    There has been a recent increasing demand to digitize Arabic books and documents, due to the fact that digital books do not lose quality over time, and can be easily sustained. Meanwhile, the number of Arabic-speaking Internet users is increasing. We propose AreCAPTCHA, a system that digitizes Arabic text by outsourcing it to native Arabic speakers, while offering protective measures to online web forms of Arabic websites. As users interact with AreCAPTCHA, we collect possible digitizations of words that were not recognized by OCR programs. We explain how the system works, the challenges we faced, and promising preliminary evaluation results.
  • Keywords
    Web sites; document image processing; natural language processing; optical character recognition; security of data; Arabic Web sites; Arabic book; Arabic document; Arabic text digitization; Arabic-speaking Internet user; AreCAPTCHA; OCR program; digital book; native Arabic speaker; online Web form; protective measure; CAPTCHAs; Databases; Educational institutions; Engines; Internet; Libraries; Optical character recognition software; Arabic; CAPTCHA; Digitization; Human computation; words recognition;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Document Analysis Systems (DAS), 2014 11th IAPR International Workshop on
  • Conference_Location
    Tours
  • Print_ISBN
    978-1-4799-3243-6
  • Type

    conf

  • DOI
    10.1109/DAS.2014.50
  • Filename
    6831018