MASTER: Multi-Aspect Non-local Network for Scene Text Recognition
arXiv:1910.02562 · doi:10.1016/j.patcog.2021.107980
Abstract
Attention-based scene text recognizers have gained huge success, which leverages a more compact intermediate representation to learn 1d- or 2d- attention by a RNN-based encoder-decoder architecture. However, such methods suffer from attention-drift problem because high similarity among encoded features leads to attention confusion under the RNN-based local attention mechanism. Moreover, RNN-based methods have low efficiency due to poor parallelization. To overcome these problems, we propose the MASTER, a self-attention based scene text recognizer that (1) not only encodes the input-output attention but also learns self-attention which encodes feature-feature and target-target relationships inside the encoder and decoder and (2) learns a more powerful and robust intermediate representation to spatial distortion, and (3) owns a great training efficiency because of high training parallelization and a high-speed inference because of an efficient memory-cache mechanism. Extensive experiments on various benchmarks demonstrate the superior performance of our MASTER on both regular and irregular scene text. Pytorch code can be found at https://github.com/wenwenyu/MASTER-pytorch, and Tensorflow code can be found at https://github.com/jiangxiluning/MASTER-TF.
Accepted by Pattern Recognition. Ning Lu and Wenwen Yu are co-first authors
References in corpus (13)
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Convolutional Sequence to Sequence Learning
- Squeeze-and-Excitation Networks
- Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition
- Focusing Attention: Towards Accurate Text Recognition in Natural Images
- Layer Normalization
- Recent Advances in Convolutional Neural Networks
- Scene Text Detection and Recognition: The Deep Learning Era
- Enhancing Energy Minimization Framework for Scene Text Recognition with Top-Down Cues
- A pooling based scene text proposal technique for scene text reading in the wild
- Symmetry-constrained Rectification Network for Scene Text Recognition
- TextProposals: a Text-specific Selective Search Algorithm for Word Spotting in the Wild
- Mask TextSpotter: An End-to-End Trainable Neural Network for Spotting Text with Arbitrary Shapes
Cited by in corpus (16)
- Aligning Correlation Information for Domain Adaptation in Action Recognition
- PingAn-VCGroup's Solution for ICDAR 2021 Competition on Scientific Literature Parsing Task B: Table Recognition to HTML
- Pure Transformer with Integrated Experts for Scene Text Recognition
- Rethinking Text Line Recognition Models
- Comprehensive Benchmark Datasets for Amharic Scene Text Detection and Recognition
- TRIG: Transformer-Based Text Recognizer with Initial Embedding Guidance
- SCATTER: Selective Context Attentional Scene Text Recognizer
- PingAn-VCGroup's Solution for ICDAR 2021 Competition on Scientific Table Image Recognition to Latex
- Hamming OCR: A Locality Sensitive Hashing Neural Network for Scene Text Recognition
- TC-OCR: TableCraft OCR for Efficient Detection & Recognition of Table Structure & Content
- SPRINT: Script-agnostic Structure Recognition in Tables
- Revisiting Classification Perspective on Scene Text Recognition
- Representation and Correlation Enhanced Encoder-Decoder Framework for Scene Text Recognition
- Utilizing Resource-Rich Language Datasets for End-to-End Scene Text Recognition in Resource-Poor Languages
- SELECT: Detecting Label Errors in Real-world Scene Text Data
- Portmanteauing Features for Scene Text Recognition