UnrealText: Synthesizing Realistic Scene Text Images from the Unreal World
arXiv:2003.10608
Abstract
Synthetic data has been a critical tool for training scene text detection and recognition models. On the one hand, synthetic word images have proven to be a successful substitute for real images in training scene text recognizers. On the other hand, however, scene text detectors still heavily rely on a large amount of manually annotated real-world images, which are expensive. In this paper, we introduce UnrealText, an efficient image synthesis method that renders realistic images via a 3D graphics engine. 3D synthetic engine provides realistic appearance by rendering scene and text as a whole, and allows for better text region proposals with access to precise scene information, e.g. normal and even object meshes. The comprehensive experiments verify its effectiveness on both scene text detection and recognition. We also generate a multilingual version for future research into multilingual scene text detection and recognition. Additionally, we re-annotate scene text recognition datasets in a case-sensitive way and include punctuation marks for more comprehensive evaluations. The code and the generated datasets are released at: https://github.com/Jyouhou/UnrealText/ .
adding experiments with Mask-RCNN
References in corpus (7)
- Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition
- R2CNN: Rotational Region CNN for Orientation Robust Scene Text Detection
- SceneNet RGB-D: 5M Photorealistic Images of Synthetic Indoor Trajectories with Ground Truth
- 2D Attentional Irregular Scene Text Recognizer
- Adversarial Generation of Training Examples: Applications to Moving Vehicle License Plate Recognition
- An Annotation Saved is an Annotation Earned: Using Fully Synthetic Training for Object Instance Detection
- Rethinking Irregular Scene Text Recognition
Cited by in corpus (9)
- Industrial Scene Text Detection with Refined Feature-attentive Network
- Text Recognition in the Wild: A Survey
- Open Images V5 Text Annotation and Yet Another Mask Text Spotter
- What If We Only Use Real Datasets for Scene Text Recognition? Toward Scene Text Recognition With Fewer Labels
- Synthetic-to-Real Unsupervised Domain Adaptation for Scene Text Detection in the Wild
- Synthesis in Style: Semantic Segmentation of Historical Documents using Synthetic Data
- OmniPrint: A Configurable Printed Character Synthesizer
- Utilizing Resource-Rich Language Datasets for End-to-End Scene Text Recognition in Resource-Poor Languages
- SDL: New data generation tools for full-level annotated document layout