Focusing Attention: Towards Accurate Text Recognition in Natural Images
arXiv:1709.02054 · doi:10.1109/ICCV.2017.543
Abstract
Scene text recognition has been a hot research topic in computer vision due to its various applications. The state of the art is the attention-based encoder-decoder framework that learns the mapping between input images and output sequences in a purely data-driven way. However, we observe that existing attention-based methods perform poorly on complicated and/or low-quality images. One major reason is that existing methods cannot get accurate alignments between feature areas and targets for such images. We call this phenomenon "attention drift". To tackle this problem, in this paper we propose the FAN (the abbreviation of Focusing Attention Network) method that employs a focusing attention mechanism to automatically draw back the drifted attention. FAN consists of two major components: an attention network (AN) that is responsible for recognizing character targets as in the existing methods, and a focusing network (FN) that is responsible for adjusting attention by evaluating whether AN pays attention properly on the target areas in the images. Furthermore, different from the existing methods, we adopt a ResNet-based network to enrich deep representations of scene text images. Extensive experiments on various benchmarks, including the IIIT5k, SVT and ICDAR datasets, show that the FAN method substantially outperforms the existing methods.
Revise the description of IC15 datasets (1811 samples)
References in corpus (6)
- Sequence to Sequence Learning with Neural Networks
- ADADELTA: An Adaptive Learning Rate Method
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Attention-Based Models for Speech Recognition
- Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition
- Joint CTC-Attention based End-to-End Speech Recognition using Multi-task Learning
Cited by in corpus (64)
- Convolutional Neural Networks with Gated Recurrent Connections
- Recognition of Handwritten Chinese Text by Segmentation: A Segment-annotation-free Approach
- CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model
- 2D Attentional Irregular Scene Text Recognizer
- TextSR: Content-Aware Text Super-Resolution Guided by Recognition
- 2D-CTC for Scene Text Recognition
- Text Detection and Recognition in the Wild: A Review
- Read Like Humans: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text Recognition
- SEED: Semantics Enhanced Encoder-Decoder Framework for Scene Text Recognition
- CDistNet: Perceiving Multi-Domain Character Distance for Robust Text Recognition
- Word-level Sign Language Recognition with Multi-stream Neural Networks Focusing on Local Regions and Skeletal Information
- RobustScanner: Dynamically Enhancing Positional Clues for Robust Text Recognition
- TextScanner: Reading Characters in Order for Robust Scene Text Recognition
- A Multi-Object Rectified Attention Network for Scene Text Recognition
- TRIG: Transformer-Based Text Recognizer with Initial Embedding Guidance
- Recurrent Calibration Network for Irregular Text Recognition
- Convolutional Character Networks
- Offline Handwritten Chinese Text Recognition with Convolutional Neural Networks
- From Two to One: A New Scene Text Recognizer with Visual Language Modeling Network
- Bidirectional Scene Text Recognition with a Single Decoder
- Handwritten Mathematical Expression Recognition with Bidirectionally Trained Transformer
- Decoupled Attention Network for Text Recognition
- Self-Supervised Learning for Text Recognition: A Critical Survey
- Improving Attention-Based Handwritten Mathematical Expression Recognition with Scale Augmentation and Drop Attention
- Scene Text Image Super-Resolution in the Wild
- Gaussian Constrained Attention Network for Scene Text Recognition
- PERT: A Progressively Region-based Network for Scene Text Removal
- Hamming OCR: A Locality Sensitive Hashing Neural Network for Scene Text Recognition
- STRIDE : Scene Text Recognition In-Device
- Primitive Representation Learning for Scene Text Recognition
- GA-DAN: Geometry-Aware Domain Adaptation Network for Scene Text Detection and Recognition
- SAFL: A Self-Attention Scene Text Recognizer with Focal Loss
- Revisiting Classification Perspective on Scene Text Recognition
- Joint Visual Semantic Reasoning: Multi-Stage Decoder for Text Recognition
- Spatial Context-based Self-Supervised Learning for Handwritten Text Recognition
- A Hybrid Vision Transformer Approach for Mathematical Expression Recognition
- Character Region Attention For Text Spotting
- SNIDER: Single Noisy Image Denoising and Rectification for Improving License Plate Recognition
- MetaHTR: Towards Writer-Adaptive Handwritten Text Recognition
- GTC: Guided Training of CTC Towards Efficient and Accurate Scene Text Recognition
- Decoupling Visual-Semantic Feature Learning for Robust Scene Text Recognition
- Towards the Unseen: Iterative Text Recognition by Distilling from Errors
- Text is Text, No Matter What: Unifying Text Recognition using Knowledge Distillation
- Representation and Correlation Enhanced Encoder-Decoder Framework for Scene Text Recognition
- Parallel Scale-wise Attention Network for Effective Scene Text Recognition
- On Vocabulary Reliance in Scene Text Recognition
- I2C2W: Image-to-Character-to-Word Transformers for Accurate Scene Text Recognition
- AE TextSpotter: Learning Visual and Linguistic Representation for Ambiguous Text Spotting
- Object-QA: Towards High Reliable Object Quality Assessment
- ReADS: A Rectified Attentional Double Supervised Network for Scene Text Recognition
- Towards Fully Automated Manga Translation
- PIMNet: A Parallel, Iterative and Mimicking Network for Scene Text Recognition
- Improving Structured Text Recognition with Regular Expression Biasing
- SAFE: Scale Aware Feature Encoder for Scene Text Recognition
- Context-Free TextSpotter for Real-Time and Mobile End-to-End Text Detection and Recognition
- IFR: Iterative Fusion Based Recognizer For Low Quality Scene Text Recognition
- Scene Text recognition with Full Normalization
- Text Recognition in Real Scenarios with a Few Labeled Samples
- Using Human Psychophysics to Evaluate Generalization in Scene Text Recognition Models
- Exploring Font-independent Features for Scene Text Recognition
- Reciprocal Feature Learning via Explicit and Implicit Tasks in Scene Text Recognition
- Scene Text Recognition with Temporal Convolutional Encoder
- Scene Text Recognition With Finer Grid Rectification
- Stratified Domain Adaptation: A Progressive Self-Training Approach for Scene Text Recognition