COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images
arXiv:1601.07140
Abstract
This paper describes the COCO-Text dataset. In recent years large-scale datasets like SUN and Imagenet drove the advancement of scene understanding and object recognition. The goal of COCO-Text is to advance state-of-the-art in text detection and recognition in natural images. The dataset is based on the MS COCO dataset, which contains images of complex everyday scenes. The images were not collected with text in mind and thus contain a broad variety of text instances. To reflect the diversity of text in natural scenes, we annotate text with (a) location in terms of a bounding box, (b) fine-grained classification into machine printed text and handwritten text, (c) classification into legible and illegible text, (d) script of the text and (e) transcriptions of legible text. The dataset contains over 173k text annotations in over 63k images. We provide a statistical analysis of the accuracy of our annotations. In addition, we present an analysis of three leading state-of-the-art photo Optical Character Recognition (OCR) approaches on our dataset. While scene text detection and recognition enjoys strong advances in recent years, we identify significant shortcomings motivating future work.
References in corpus (3)
Cited by in corpus (76)
- TextBoxes++: A Single-Shot Oriented Scene Text Detector
- Object Detection in 20 Years: A Survey
- Rosetta: Large scale system for text detection and recognition in images
- Scene Text Detection via Holistic, Multi-Channel Prediction
- Detecting Curve Text in the Wild: New Dataset and New Solution
- OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
- EAST: An Efficient and Accurate Scene Text Detector
- Recent Advances in Object Detection in the Age of Deep Convolutional Neural Networks
- DeepText: A Unified Framework for Text Proposal Generation and Text Detection in Natural Images
- What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model Analysis
- Weak Supervision for Generating Pixel-Level Annotations in Scene Text Segmentation
- Recursive Recurrent Nets with Attention Modeling for OCR in the Wild
- Selective Feature Connection Mechanism: Concatenating Multi-layer CNN Features with a Feature Selector
- Single Shot Text Detector with Regional Attention
- PixelLink: Detecting Scene Text via Instance Segmentation
- Text Recognition in the Wild: A Survey
- Mask TextSpotter: An End-to-End Trainable Neural Network for Spotting Text with Arbitrary Shapes
- Detecting Multi-Oriented Text with Corner-based Region Proposals
- Show, Attend and Read: A Simple and Strong Baseline for Irregular Text Recognition
- WordSup: Exploiting Word Annotations for Character based Text Detection
- Multi-Oriented Scene Text Detection via Corner Localization and Region Segmentation
- ABCNet: Real-time Scene Text Spotting with Adaptive Bezier-Curve Network
- Stroke-Based Scene Text Erasing Using Synthetic Data for Training
- Character-Based Handwritten Text Transcription with Attention Networks
- ICDAR2019 Robust Reading Challenge on Arbitrary-Shaped Text (RRC-ArT)
- Class-incremental Learning via Deep Model Consolidation
- Text Detection and Recognition in the Wild: A Review
- PAN++: Towards Efficient and Accurate End-to-End Spotting of Arbitrarily-Shaped Text
- ABCNet v2: Adaptive Bezier-Curve Network for Real-time End-to-end Text Spotting
- TAP: Text-Aware Pre-training for Text-VQA and Text-Caption
- Rotation-Sensitive Regression for Oriented Scene Text Detection
- Iterative Answer Prediction with Pointer-Augmented Multimodal Transformers for TextVQA
- WeText: Scene Text Detection under Weak Supervision
- Structured Multimodal Attentions for TextVQA
- On the General Value of Evidence, and Bilingual Scene-Text Visual Question Answering
- ICDAR 2019 Competition on Large-scale Street View Text with Partial Labeling -- RRC-LSVT
- A Binary Convolutional Encoder-decoder Network for Real-time Natural Scene Text Processing
- WordFence: Text Detection in Natural Images with Border Awareness
- Verisimilar Image Synthesis for Accurate Detection and Recognition of Texts in Scenes
- Bidirectional Scene Text Recognition with a Single Decoder
- AdaDNNs: Adaptive Ensemble of Deep Neural Networks for Scene Text Recognition
- Towards End-to-End Text Spotting in Natural Scenes
- EfficientCLIP: Efficient Cross-Modal Pre-training by Ensemble Confident Learning and Language Modeling
- Chinese Street View Text: Large-scale Chinese Text Reading with Partially Supervised Learning
- FC2RN: A Fully Convolutional Corner Refinement Network for Accurate Multi-Oriented Scene Text Detection
- What If We Only Use Real Datasets for Scene Text Recognition? Toward Scene Text Recognition With Fewer Labels
- You Only Recognize Once: Towards Fast Video Text Spotting
- All You Need Is Boundary: Toward Arbitrary-Shaped Text Spotting
- DeRPN: Taking a further step toward more general object detection
- An efficient and perceptually motivated auditory neural encoding and decoding algorithm for spiking neural networks
- RRPN++: Guidance Towards More Accurate Scene Text Detection
- Traditional Chinese Synthetic Datasets Verified with Labeled Data for Scene Text Recognition
- Self-Guiding Multimodal LSTM - when we do not have a perfect training dataset for image captioning
- Exploring the Capacity of an Orderless Box Discretization Network for Multi-orientation Scene Text Detection
- A New Unified Method for Detecting Text from Marathon Runners and Sports Players in Video
- FedOCR: Communication-Efficient Federated Learning for Scene Text Recognition
- TextTubes for Detecting Curved Text in the Wild
- Rethinking Text Segmentation: A Novel Dataset and A Text-Specific Refinement Approach
- TextSLAM: Visual SLAM with Planar Text Features
- Scale-Invariant Multi-Oriented Text Detection in Wild Scene Images
- Towards Spatio-Temporal Video Scene Text Detection via Temporal Clustering
- Scene Text Retrieval via Joint Text Detection and Similarity Learning
- Scene Text Detection with Selected Anchor
- End-to-End Interpretation of the French Street Name Signs Dataset
- Single Shot Scene Text Retrieval
- Advances of Scene Text Datasets
- Localize, Group, and Select: Boosting Text-VQA by Scene Text Modeling
- Boosting Image Recognition with Non-differentiable Constraints
- Semantic Relatedness Based Re-ranker for Text Spotting
- Input Fast-Forwarding for Better Deep Learning
- Challenging Images For Minds and Machines
- A Multiplexed Network for End-to-End, Multilingual OCR
- Exploring Font-independent Features for Scene Text Recognition
- A method for detecting text of arbitrary shapes in natural scenes that improves text spotting
- An Image Dataset of Text Patches in Everyday Scenes
- ICDAR 2021 Competition on Document VisualQuestion Answering