Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition
arXiv:1406.2227
Abstract
In this work we present a framework for the recognition of natural scene text. Our framework does not require any human-labelled data, and performs word recognition on the whole image holistically, departing from the character based recognition systems of the past. The deep neural network models at the centre of this framework are trained solely on data produced by a synthetic text generation engine -- synthetic data that is highly realistic and sufficient to replace real data, giving us infinite amounts of training data. This excess of data exposes new possibilities for word recognition models, and here we consider three models, each one "reading" words in a different way: via 90k-way dictionary encoding, character sequence encoding, and bag-of-N-grams encoding. In the scenarios of language based and completely unconstrained text recognition we greatly improve upon state-of-the-art performance on standard datasets, using our fast, simple machinery and requiring zero data-acquisition costs.
References in corpus (1)
Cited by in corpus (86)
- Multiple Object Recognition with Visual Attention
- Focusing Attention: Towards Accurate Text Recognition in Natural Images
- Rosetta: Large scale system for text detection and recognition in images
- Scene Text Detection via Holistic, Multi-Channel Prediction
- Deep Structured Output Learning for Unconstrained Text Recognition
- Are Deepfakes Concerning? Analyzing Conversations of Deepfakes on Reddit and Exploring Societal Implications
- Scene Text Recognition with Sliding Convolutional Character Models
- Convolutional Neural Networks with Gated Recurrent Connections
- Reading Scene Text with Attention Convolutional Sequence Modeling
- 2D Attentional Irregular Scene Text Recognizer
- TextSR: Content-Aware Text Super-Resolution Guided by Recognition
- WordSup: Exploiting Word Annotations for Character based Text Detection
- Adversarial Generation of Training Examples: Applications to Moving Vehicle License Plate Recognition
- 2D-CTC for Scene Text Recognition
- A Survey of Deep Learning Approaches for OCR and Document Understanding
- Text Detection and Recognition in the Wild: A Review
- Read Like Humans: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text Recognition
- Rethinking Text Line Recognition Models
- Efficient Parameter-free Clustering Using First Neighbor Relations
- Transfer Learning for Scene Text Recognition in Indian Languages
- ResearchDoom and CocoDoom: Learning Computer Vision with Games
- RobustScanner: Dynamically Enhancing Positional Clues for Robust Text Recognition
- A Binary Convolutional Encoder-decoder Network for Real-time Natural Scene Text Processing
- TextScanner: Reading Characters in Order for Robust Scene Text Recognition
- A Multi-Object Rectified Attention Network for Scene Text Recognition
- WordFence: Text Detection in Natural Images with Border Awareness
- Supervised mid-level features for word image representation
- Visual-Semantic Transformer for Scene Text Recognition
- A Survey on Deep Domain Adaptation and Tiny Object Detection Challenges, Techniques and Datasets
- The End-of-End-to-End: A Video Understanding Pentathlon Challenge (2020)
- From Two to One: A New Scene Text Recognizer with Visual Language Modeling Network
- Recurrent Calibration Network for Irregular Text Recognition
- FOTS: Fast Oriented Text Spotting with a Unified Network
- Decoupled Attention Network for Text Recognition
- KISS: Keeping It Simple for Scene Text Recognition
- Scene Text Image Super-Resolution in the Wild
- Smart Library: Identifying Books in a Library using Richly Supervised Deep Scene Text Reading
- AdaDNNs: Adaptive Ensemble of Deep Neural Networks for Scene Text Recognition
- Towards Boosting the Accuracy of Non-Latin Scene Text Recognition
- Interpreting Neural Networks Using Flip Points
- Hamming OCR: A Locality Sensitive Hashing Neural Network for Scene Text Recognition
- GA-DAN: Geometry-Aware Domain Adaptation Network for Scene Text Detection and Recognition
- Human-Expert-Level Brain Tumor Detection Using Deep Learning with Data Distillation and Augmentation
- Navigation Agents for the Visually Impaired: A Sidewalk Simulator and Experiments
- SAFL: A Self-Attention Scene Text Recognizer with Focal Loss
- Improving Performance of End-to-End ASR on Numeric Sequences
- Why You Should Try the Real Data for the Scene Text Recognition
- Joint Visual Semantic Reasoning: Multi-Stage Decoder for Text Recognition
- CMFN: Cross-Modal Fusion Network for Irregular Scene Text Recognition
- DropSample: A New Training Method to Enhance Deep Convolutional Neural Networks for Large-Scale Unconstrained Handwritten Chinese Character Recognition
- A New Perspective for Flexible Feature Gathering in Scene Text Recognition Via Character Anchor Pooling
- Context Augmentation for Convolutional Neural Networks
- Tracking e-cigarette warning label compliance on Instagram with deep learning
- Large Scale Font Independent Urdu Text Recognition System
- Parallel Scale-wise Attention Network for Effective Scene Text Recognition
- A study of the effect of the illumination model on the generation of synthetic training datasets
- Rethinking Text Segmentation: A Novel Dataset and A Text-Specific Refinement Approach
- Text is Text, No Matter What: Unifying Text Recognition using Knowledge Distillation
- Towards the Unseen: Iterative Text Recognition by Distilling from Errors
- OmniPrint: A Configurable Printed Character Synthesizer
- Integrating Scene Text and Visual Appearance for Fine-Grained Image Classification
- On Vocabulary Reliance in Scene Text Recognition
- ReADS: A Rectified Attentional Double Supervised Network for Scene Text Recognition
- Watermark retrieval from 3D printed objects via synthetic data training
- Adaptive Text Recognition through Visual Matching
- Sequence to sequence learning for unconstrained scene text recognition
- Rudder: A Cross Lingual Video and Text Retrieval Dataset
- End-to-End Interpretation of the French Street Name Signs Dataset
- Towards Fully Automated Manga Translation
- Robust Handwriting Recognition with Limited and Noisy Data
- Scene Text Recognition With Finer Grid Rectification
- A Machine Learning Framework for Data Ingestion in Document Images
- Implicit Feature Alignment: Learn to Convert Text Recognizer to Text Spotter
- Labeled Data Generation with Inexact Supervision
- SAFE: Scale Aware Feature Encoder for Scene Text Recognition
- Utilizing High-level Visual Feature for Indoor Shopping Mall Navigation
- Text Recognition in Real Scenarios with a Few Labeled Samples
- Portmanteauing Features for Scene Text Recognition
- MailLeak: Obfuscation-Robust Character Extraction Using Transfer Learning
- Improving Word Recognition using Multiple Hypotheses and Deep Embeddings
- Distantly Supervised Semantic Text Detection and Recognition for Broadcast Sports Videos Understanding
- Scene Text recognition with Full Normalization
- Meta Self-Learning for Multi-Source Domain Adaptation: A Benchmark
- IFR: Iterative Fusion Based Recognizer For Low Quality Scene Text Recognition
- Learning to Infer User Interface Attributes from Images
- SmartTennisTV: Automatic indexing of tennis videos