Text-Attentional Convolutional Neural Networks for Scene Text Detection
arXiv:1510.03283 · doi:10.1109/TIP.2016.2547588
Abstract
Recent deep learning models have demonstrated strong capabilities for classifying text and non-text components in natural images. They extract a high-level feature computed globally from a whole image component (patch), where the cluttered background information may dominate true text features in the deep representation. This leads to less discriminative power and poorer robustness. In this work, we present a new system for scene text detection by proposing a novel Text-Attentional Convolutional Neural Network (Text-CNN) that particularly focuses on extracting text-related regions and features from the image components. We develop a new learning mechanism to train the Text-CNN with multi-level and rich supervised information, including text region mask, character label, and binary text/nontext information. The rich supervision information enables the Text-CNN with a strong capability for discriminating ambiguous texts, and also increases its robustness against complicated background components. The training process is formulated as a multi-task learning problem, where low-level supervised information greatly facilitates main task of text/non-text classification. In addition, a powerful low-level detector called Contrast- Enhancement Maximally Stable Extremal Regions (CE-MSERs) is developed, which extends the widely-used MSERs by enhancing intensity contrast between text patterns and background. This allows it to detect highly challenging text patterns, resulting in a higher recall. Our approach achieved promising results on the ICDAR 2013 dataset, with a F-measure of 0.82, improving the state-of-the-art results substantially.
To appear in IEEE Trans. on Image Processing, 2016
References in corpus (4)
Cited by in corpus (34)
- Arbitrary-Oriented Scene Text Detection via Rotation Proposals
- TextBoxes++: A Single-Shot Oriented Scene Text Detector
- Object Detection in 20 Years: A Survey
- TextField: Learning A Deep Direction Field for Irregular Scene Text Detection
- An Energy-Efficient FPGA-based Deconvolutional Neural Networks Accelerator for Single Image Super-Resolution
- Locally-Supervised Deep Hybrid Model for Scene Recognition
- Accurate Text Localization in Natural Image with Cascaded Convolutional Text Network
- Single Shot Text Detector with Regional Attention
- Deep Direct Regression for Multi-Oriented Scene Text Detection
- Real-time Scene Text Detection with Differentiable Binarization
- Multi-Oriented Text Detection and Verification in Video Frames and Scene Images
- WordSup: Exploiting Word Annotations for Character based Text Detection
- VWA: Hardware Efficient Vectorwise Accelerator for Convolutional Neural Network
- Reading Scene Text in Deep Convolutional Sequences
- IncepText: A New Inception-Text Module with Deformable PSROI Pooling for Multi-Oriented Scene Text Detection
- Scene Text Synthesis for Efficient and Effective Deep Network Training
- A Light Dual-Task Neural Network for Haze Removal
- A pooling based scene text proposal technique for scene text reading in the wild
- Detecting Text in Natural Image with Connectionist Text Proposal Network
- WeText: Scene Text Detection under Weak Supervision
- Verisimilar Image Synthesis for Accurate Detection and Recognition of Texts in Scenes
- Mask TextSpotter v3: Segmentation Proposal Network for Robust Scene Text Spotting
- VSCNN: Convolution Neural Network Accelerator With Vector Sparsity
- Urdu text in natural scene images: a new dataset and preliminary text detection
- Deep Neural Network for Semantic-based Text Recognition in Images
- GA-DAN: Geometry-Aware Domain Adaptation Network for Scene Text Detection and Recognition
- SA-CNN: Application to text categorization issues using simulated annealing-based convolutional neural network optimization
- Cascaded Segmentation-Detection Networks for Word-Level Text Spotting
- Exploring the Capacity of an Orderless Box Discretization Network for Multi-orientation Scene Text Detection
- MT3S: Mobile Turkish Scene Text-to-Speech System for the Visually Impaired
- Detecting Text in the Wild with Deep Character Embedding Network
- Learning Robust Feature Representations for Scene Text Detection
- Accurate Scene Text Detection through Border Semantics Awareness and Bootstrapping
- Scene Text recognition with Full Normalization