• DocumentCode
    157911
  • Title

    Video text detection and recognition: Dataset and benchmark

  • Author

    Phuc Xuan Nguyen ; Kai Wang ; Belongie, Serge

  • Author_Institution
    Dept. of Comput. Sci. & Eng., Univ. of California San Diego, La Jolla, CA, USA
  • fYear
    2014
  • fDate
    24-26 March 2014
  • Firstpage
    776
  • Lastpage
    783
  • Abstract
    This paper focuses on the problem of text detection and recognition in videos. Even though text detection and recognition in images has seen much progress in recent years, relatively little work has been done to extend these solutions to the video domain. In this work, we extend an existing end-to-end solution for text recognition in natural images to video. We explore a variety of methods for training local character models and explore methods to capitalize on the temporal redundancy of text in video. We present detection performance using the Video Analysis and Content Extraction (VACE) benchmarking framework on the ICDAR 2013 Robust Reading Challenge 3 video dataset and on a new video text dataset. We also propose a new performance metric based on precision-recall curves to measure the performance of text recognition in videos. Using this metric, we provide early video text recognition results on the above mentioned datasets.
  • Keywords
    image recognition; learning (artificial intelligence); text analysis; video signal processing; ICDAR 2013 Robust Reading Challenge; image recognition; text recognition; video analysis and content extraction benchmarking framework; video recognition; video text detection; Benchmark testing; Measurement; Smoothing methods; Text recognition; Training; Training data; YouTube;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Applications of Computer Vision (WACV), 2014 IEEE Winter Conference on
  • Conference_Location
    Steamboat Springs, CO
  • Type

    conf

  • DOI
    10.1109/WACV.2014.6836024
  • Filename
    6836024