Scene Text Synthesis for Efficient and Effective Deep Network Training
arXiv:1901.09193
Abstract
A large amount of annotated training images is critical for training accurate and robust deep network models but the collection of a large amount of annotated training images is often time-consuming and costly. Image synthesis alleviates this constraint by generating annotated training images automatically by machines which has attracted increasing interest in the recent deep learning research. We develop an innovative image synthesis technique that composes annotated training images by realistically embedding foreground objects of interest (OOI) into background images. The proposed technique consists of two key components that in principle boost the usefulness of the synthesized images in deep network training. The first is context-aware semantic coherence which ensures that the OOI are placed around semantically coherent regions within the background image. The second is harmonious appearance adaptation which ensures that the embedded OOI are agreeable to the surrounding background from both geometry alignment and appearance realism. The proposed technique has been evaluated over two related but very different computer vision challenges, namely, scene text detection and scene text recognition. Experiments over a number of public datasets demonstrate the effectiveness of our proposed image synthesis technique - the use of our synthesized images in deep network training is capable of achieving similar or even better scene text detection and scene text recognition performance as compared with using real images.
This work has been merged into another project
References in corpus (12)
- Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition
- Focusing Attention: Towards Accurate Text Recognition in Natural Images
- Unsupervised Image-to-Image Translation Networks
- TextBoxes: A Fast Text Detector with a Single Deep Neural Network
- Scene Text Detection via Holistic, Multi-Channel Prediction
- PVANET: Deep but Lightweight Neural Networks for Real-time Object Detection
- Deep Structured Output Learning for Unconstrained Text Recognition
- Deep Matching Prior Network: Toward Tighter Multi-oriented Text Detection
- Single Shot Text Detector with Regional Attention
- ICDAR2017 Competition on Reading Chinese Text in the Wild (RCTW-17)
- Verisimilar Image Synthesis for Accurate Detection and Recognition of Texts in Scenes
- Spatial Fusion GAN for Image Synthesis
Cited by in corpus (13)
- Blind Image Super-Resolution via Contrastive Representation Learning
- ESIR: End-to-end Scene Text Recognition via Iterative Image Rectification
- EMLight: Lighting Estimation via Spherical Distribution Approximation
- Spatial Fusion GAN for Image Synthesis
- Deep Monocular 3D Human Pose Estimation via Cascaded Dimension-Lifting
- Adversarial Image Composition with Auxiliary Illumination
- FBC-GAN: Diverse and Flexible Image Synthesis via Foreground-Background Composition
- Bi-level Feature Alignment for Versatile Image Translation and Manipulation
- GA-DAN: Geometry-Aware Domain Adaptation Network for Scene Text Detection and Recognition
- Spatial-Aware GAN for Unsupervised Person Re-identification
- Hierarchy Composition GAN for High-fidelity Image Synthesis
- Scene Text recognition with Full Normalization
- A Survey on Adversarial Image Synthesis