Text Recognition in the Wild: A Survey
arXiv:2005.03492
Abstract
The history of text can be traced back over thousands of years. Rich and precise semantic information carried by text is important in a wide range of vision-based application scenarios. Therefore, text recognition in natural scenes has been an active research field in computer vision and pattern recognition. In recent years, with the rise and development of deep learning, numerous methods have shown promising in terms of innovation, practicality, and efficiency. This paper aims to (1) summarize the fundamental problems and the state-of-the-art associated with scene text recognition; (2) introduce new insights and ideas; (3) provide a comprehensive review of publicly available resources; (4) point out directions for future work. In summary, this literature review attempts to present the entire picture of the field of scene text recognition. It provides a comprehensive reference for people entering this field, and could be helpful to inspire future research. Related resources are available at our Github repository: https://github.com/HCIILAB/Scene-Text-Recognition.
Accepted by ACM Computing Surveys
References in corpus (13)
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition
- Focusing Attention: Towards Accurate Text Recognition in Natural Images
- TextBoxes: A Fast Text Detector with a Single Deep Neural Network
- Detecting Curve Text in the Wild: New Dataset and New Solution
- Deep Structured Output Learning for Unconstrained Text Recognition
- Scene Text Recognition with Sliding Convolutional Character Models
- Towards Accurate Scene Text Recognition with Semantic Reasoning Networks
- UnrealText: Synthesizing Realistic Scene Text Images from the Unreal World
- TextSR: Content-Aware Text Super-Resolution Guided by Recognition
- 2D-CTC for Scene Text Recognition
- End-to-End Information Extraction by Character-Level Embedding and Multi-Stage Attentional U-Net
- Separating Content from Style Using Adversarial Learning for Recognizing Text in the Wild