End-to-End Text Recognition with Hybrid HMM Maxout Models
arXiv:1310.1811
Abstract
The problem of detecting and recognizing text in natural scenes has proved to be more challenging than its counterpart in documents, with most of the previous work focusing on a single part of the problem. In this work, we propose new solutions to the character and word recognition problems and then show how to combine these solutions in an end-to-end text-recognition system. We do so by leveraging the recently introduced Maxout networks along with hybrid HMM models that have proven useful for voice recognition. Using these elements, we build a tunable and highly accurate recognition system that beats state-of-the-art results on all the sub-problems for both the ICDAR 2003 and SVT benchmark datasets.
9 pages, 7 figures
References in corpus (3)
Cited by in corpus (25)
- TextBoxes++: A Single-Shot Oriented Scene Text Detector
- Focusing Attention: Towards Accurate Text Recognition in Natural Images
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- TextBoxes: A Fast Text Detector with a Single Deep Neural Network
- Deep Structured Output Learning for Unconstrained Text Recognition
- Scene Text Recognition with Sliding Convolutional Character Models
- An End-to-End Trainable Neural Network for Image-based Sequence Recognition and Its Application to Scene Text Recognition
- Recursive Recurrent Nets with Attention Modeling for OCR in the Wild
- Convolutional Neural Networks with Gated Recurrent Connections
- Reading Scene Text with Attention Convolutional Sequence Modeling
- Text Recognition in the Wild: A Survey
- Reading Scene Text in Deep Convolutional Sequences
- Edit Probability for Scene Text Recognition
- Scene Text Synthesis for Efficient and Effective Deep Network Training
- 2D-CTC for Scene Text Recognition
- End-to-End Subtitle Detection and Recognition for Videos in East Asian Languages via CNN Ensemble with Near-Human-Level Performance
- Reading Text in the Wild with Convolutional Neural Networks
- NRTR: A No-Recurrence Sequence-to-Sequence Model For Scene Text Recognition
- SCAN: Sliding Convolutional Attention Network for Scene Text Recognition
- TextProposals: a Text-specific Selective Search Algorithm for Word Spotting in the Wild
- AdaDNNs: Adaptive Ensemble of Deep Neural Networks for Scene Text Recognition
- Learning to Read by Spelling: Towards Unsupervised Text Recognition
- AON: Towards Arbitrarily-Oriented Text Recognition
- Inductive Visual Localisation: Factorised Training for Superior Generalisation
- Simultaneous Recognition of Horizontal and Vertical Text in Natural Images