Rich feature hierarchies for accurate object detection and semantic segmentation
arXiv:1311.2524
Abstract
Object detection performance, as measured on the canonical PASCAL VOC dataset, has plateaued in the last few years. The best-performing methods are complex ensemble systems that typically combine multiple low-level image features with high-level context. In this paper, we propose a simple and scalable detection algorithm that improves mean average precision (mAP) by more than 30% relative to the previous best result on VOC 2012---achieving a mAP of 53.3%. Our approach combines two key insights: (1) one can apply high-capacity convolutional neural networks (CNNs) to bottom-up region proposals in order to localize and segment objects and (2) when labeled training data is scarce, supervised pre-training for an auxiliary task, followed by domain-specific fine-tuning, yields a significant performance boost. Since we combine region proposals with CNNs, we call our method R-CNN: Regions with CNN features. We also compare R-CNN to OverFeat, a recently proposed sliding-window detector based on a similar CNN architecture. We find that R-CNN outperforms OverFeat by a large margin on the 200-class ILSVRC2013 detection dataset. Source code for the complete system is available at http://www.cs.berkeley.edu/~rbg/rcnn.
Extended version of our CVPR 2014 paper; latest update (v5) includes results using deeper networks (see Appendix G. Changelog)
Cited by in corpus (295)
- Focal Loss for Dense Object Detection
- Fully Convolutional Networks for Semantic Segmentation
- Learning Deconvolution Network for Semantic Segmentation
- Training Deep Neural Networks on Noisy Labels with Bootstrapping
- Deformable Convolutional Networks
- Text Understanding from Scratch
- Fast-SCNN: Fast Semantic Segmentation Network
- VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection
- Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
- LocalViT: Analyzing Locality in Vision Transformers
- End-to-End Learning of Geometry and Context for Deep Stereo Regression
- Multi-view Convolutional Neural Networks for 3D Shape Recognition
- Visual Saliency Based on Multiscale Deep Features
- A Survey on Deep Learning-based Architectures for Semantic Segmentation on 2D images
- Unsupervised Learning of Visual Representations using Videos
- Deformable ConvNets v2: More Deformable, Better Results
- FusionNet: 3D Object Classification Using Multiple Data Representations
- DeepDriving: Learning Affordance for Direct Perception in Autonomous Driving
- An Entropy-based Pruning Method for CNN Compression
- Examining the Impact of Blur on Recognition by Convolutional Networks
- Temporal Action Detection with Structured Segment Networks
- Deep Contrast Learning for Salient Object Detection
- A Pursuit of Temporal Accuracy in General Activity Detection
- Context Encoding for Semantic Segmentation
- Deep Semantic Ranking Based Hashing for Multi-Label Image Retrieval
- R-C3D: Region Convolutional 3D Network for Temporal Activity Detection
- Optimization for Arbitrary-Oriented Object Detection via Representation Invariance Loss
- Fully Convolutional Neural Networks for Crowd Segmentation
- Adapting Mask-RCNN for Automatic Nucleus Segmentation
- Face Detection through Scale-Friendly Deep Convolutional Networks
- Single-Shot Refinement Neural Network for Object Detection
- Light-Weight RefineNet for Real-Time Semantic Segmentation
- Flow-Guided Feature Aggregation for Video Object Detection
- Deep Reinforcement Learning for Visual Object Tracking in Videos
- What Do We Understand About Convolutional Networks?
- Learning Deep Feature Representations with Domain Guided Dropout for Person Re-identification
- Rethinking ImageNet Pre-training
- Squared Earth Mover's Distance-based Loss for Training Deep Neural Networks
- DSOD: Learning Deeply Supervised Object Detectors from Scratch
- Segmentation of cell-level anomalies in electroluminescence images of photovoltaic modules
- From Captions to Visual Concepts and Back
- Perceptual Generative Adversarial Networks for Small Object Detection
- Towards Interpretable Deep Neural Networks by Leveraging Adversarial Examples
- xView: Objects in Context in Overhead Imagery
- Deep Image Retrieval: Learning global representations for image search
- Recognizing Partial Biometric Patterns
- DenseCap: Fully Convolutional Localization Networks for Dense Captioning
- How Far are We from Solving Pedestrian Detection?
- ME R-CNN: Multi-Expert R-CNN for Object Detection
- Exploiting Image-trained CNN Architectures for Unconstrained Video Classification
- Learning Multi-Domain Convolutional Neural Networks for Visual Tracking
- Complex-YOLO: Real-time 3D Object Detection on Point Clouds
- Adversarial Complementary Learning for Weakly Supervised Object Localization
- Deep Convolution Networks for Compression Artifacts Reduction
- Learning to track for spatio-temporal action localization
- Contextual Action Recognition with R*CNN
- Eye Tracking for Everyone
- Illuminating Pedestrians via Simultaneous Detection & Segmentation
- CityPersons: A Diverse Dataset for Pedestrian Detection
- An End-to-End Trainable Neural Network for Image-based Sequence Recognition and Its Application to Scene Text Recognition
- Inside-Outside Net: Detecting Objects in Context with Skip Pooling and Recurrent Neural Networks
- Learning like a Child: Fast Novel Visual Concept Learning from Sentence Descriptions of Images
- Visual Translation Embedding Network for Visual Relation Detection
- Detecting and Recognizing Human-Object Interactions
- Deformable Part Models are Convolutional Neural Networks
- Deep Matching Prior Network: Toward Tighter Multi-oriented Text Detection
- Visual Relationship Detection with Internal and External Linguistic Knowledge Distillation
- Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
- Deep Variation-structured Reinforcement Learning for Visual Relationship and Attribute Detection
- Predicting Complete 3D Models of Indoor Scenes
- Skeleton based action recognition using translation-scale invariant image mapping and multi-scale deep cnn
- Boosting Adversarial Attacks with Momentum
- Integrated Object Detection and Tracking with Tracklet-Conditioned Detection
- Accurate Single Stage Detector Using Recurrent Rolling Convolution
- SpotTune: Transfer Learning through Adaptive Fine-tuning
- Deep Discriminative Clustering Analysis
- Data Distillation: Towards Omni-Supervised Learning
- Mobile Video Object Detection with Temporally-Aware Feature Maps
- Multi-Channel CNN-based Object Detection for Enhanced Situation Awareness
- Unsupervised Object Discovery and Localization in the Wild: Part-based Matching with Bottom-up Region Proposals
- Two-Phase Learning for Weakly Supervised Object Localization
- ReconNet: Non-Iterative Reconstruction of Images from Compressively Sensed Random Measurements
- Unsupervised feature learning by augmenting single images
- Automated crater shape retrieval using weakly-supervised deep learning
- Tracking Randomly Moving Objects on Edge Box Proposals
- Beyond Local Search: Tracking Objects Everywhere with Instance-Specific Proposals
- Weakly- and Semi-Supervised Object Detection with Expectation-Maximization Algorithm
- Learning non-maximum suppression
- Fast object detection in compressed JPEG Images
- Piggyback: Adapting a Single Network to Multiple Tasks by Learning to Mask Weights
- Deep Convolutional Neural Networks with Merge-and-Run Mappings
- Monocular Object Instance Segmentation and Depth Ordering with CNNs
- Unconstrained Facial Landmark Localization with Backbone-Branches Fully-Convolutional Networks
- On Learning to Think: Algorithmic Information Theory for Novel Combinations of Reinforcement Learning Controllers and Recurrent Neural World Models
- Rethinking the Faster R-CNN Architecture for Temporal Action Localization
- Dice Loss for Data-imbalanced NLP Tasks
- Deep Region Hashing for Efficient Large-scale Instance Search from Images
- TableBank: A Benchmark Dataset for Table Detection and Recognition
- Diversify and Match: A Domain Adaptive Representation Learning Paradigm for Object Detection
- DeepMVS: Learning Multi-view Stereopsis
- Inferring 3D Object Pose in RGB-D Images
- Designing Deep Networks for Surface Normal Estimation
- DeepBox: Learning Objectness with Convolutional Networks
- A Discriminative CNN Video Representation for Event Detection
- Scene Graph Generation from Objects, Phrases and Region Captions
- Impression Network for Video Object Detection
- End-to-end people detection in crowded scenes
- Fusing Multi-Stream Deep Networks for Video Classification
- MegDet: A Large Mini-Batch Object Detector
- Deep Pyramidal Residual Networks
- Joint Object and Part Segmentation using Deep Learned Potentials
- Viewpoints and Keypoints
- Natural Language Object Retrieval
- Effect of Annotation Errors on Drone Detection with YOLOv3
- Learning to Detect Human-Object Interactions
- MITOS-RCNN: A Novel Approach to Mitotic Figure Detection in Breast Cancer Histopathology Images using Region Based Convolutional Neural Networks
- Beyond Frontal Faces: Improving Person Recognition Using Multiple Cues
- Feature Agglomeration Networks for Single Stage Face Detection
- A semi-supervised self-training method to develop assistive intelligence for segmenting multiclass bridge elements from inspection videos
- Structure Inference Net: Object Detection Using Scene-Level Context and Instance-Level Relationships
- Soft Proposal Networks for Weakly Supervised Object Localization
- Fine-grained Categorization and Dataset Bootstrapping using Deep Metric Learning with Humans in the Loop
- Multi-Instance Visual-Semantic Embedding
- Not All Pixels Are Equal: Difficulty-aware Semantic Segmentation via Deep Layer Cascade
- 3D Bounding Box Estimation Using Deep Learning and Geometry
- Deep Learning for Single-View Instance Recognition
- Repulsion Loss: Detecting Pedestrians in a Crowd
- Attentional Network for Visual Object Detection
- Material Recognition in the Wild with the Materials in Context Database
- Instance-Level Segmentation for Autonomous Driving with Deep Densely Connected MRFs
- Learning Convolutional Networks for Content-weighted Image Compression
- Weakly-supervised learning of visual relations
- Deep Learning for Object Saliency Detection and Image Segmentation
- Deep Cuboid Detection: Beyond 2D Bounding Boxes
- Iterative Instance Segmentation
- Learning Discriminative Motion Features Through Detection
- Improved Person Detection on Omnidirectional Images with Non-maxima Suppression
- Dense Optical Flow Prediction from a Static Image
- Improving Interpretability of Deep Neural Networks with Semantic Information
- Visualizing and Understanding Deep Texture Representations
- Deep learning for class-generic object detection
- Doubly Attentive Transformer Machine Translation
- Borrowing Treasures from the Wealthy: Deep Transfer Learning through Selective Joint Fine-tuning
- Breaking Batch Normalization for better explainability of Deep Neural Networks through Layer-wise Relevance Propagation
- Solution for Large-Scale Hierarchical Object Detection Datasets with Incomplete Annotation and Data Imbalance
- Convolutional Neural Networks for joint object detection and pose estimation: A comparative study
- EraseReLU: A Simple Way to Ease the Training of Deep Convolution Neural Networks
- Automatic Concept Discovery from Parallel Text and Visual Corpora
- Training object class detectors with click supervision
- A New Convolutional Network-in-Network Structure and Its Applications in Skin Detection, Semantic Segmentation, and Artifact Reduction
- Object Detection Free Instance Segmentation With Labeling Transformations
- Learning to Segment Every Thing
- Pose from Action: Unsupervised Learning of Pose Features based on Motion
- Mining Discriminative Triplets of Patches for Fine-Grained Classification
- Combining the Best of Graphical Models and ConvNets for Semantic Segmentation
- Single-Shot Object Detection with Enriched Semantics
- Detecting Oriented Text in Natural Images by Linking Segments
- Digging Deep into the layers of CNNs: In Search of How CNNs Achieve View Invariance
- DeepEdge: A Multi-Scale Bifurcated Deep Network for Top-Down Contour Detection
- Deep Learning of Appearance Models for Online Object Tracking
- LCNN: Lookup-based Convolutional Neural Network
- Control Distance IoU and Control Distance IoU Loss Function for Better Bounding Box Regression
- Multi-scale Location-aware Kernel Representation for Object Detection
- Generic Object Detection With Dense Neural Patterns and Regionlets
- CNN Based Hashing for Image Retrieval
- TAFE-Net: Task-Aware Feature Embeddings for Low Shot Learning
- Situational Object Boundary Detection
- Learning Instance Occlusion for Panoptic Segmentation
- Pyramid R-CNN: Towards Better Performance and Adaptability for 3D Object Detection
- VrR-VG: Refocusing Visually-Relevant Relationships
- Iterative Deep Learning for Network Topology Extraction
- Optimizing Video Object Detection via a Scale-Time Lattice
- The Role of Context Selection in Object Detection
- Single-Shot Bidirectional Pyramid Networks for High-Quality Object Detection
- Pointwise Convolutional Neural Networks
- Weakly- and Self-Supervised Learning for Content-Aware Deep Image Retargeting
- An End-to-End Approach to Natural Language Object Retrieval via Context-Aware Deep Reinforcement Learning
- Reversible Recursive Instance-level Object Segmentation
- Convolutional Models for Joint Object Categorization and Pose Estimation
- Actions and Attributes from Wholes and Parts
- Learning Sparse High Dimensional Filters: Image Filtering, Dense CRFs and Bilateral Neural Networks
- Turning a Blind Eye: Explicit Removal of Biases and Variation from Deep Neural Network Embeddings
- Learning Visual Features from Large Weakly Supervised Data
- Amodal Completion and Size Constancy in Natural Scenes
- Differential Angular Imaging for Material Recognition
- Real-Time End-to-End Action Detection with Two-Stream Networks
- Deep Learning for Image Denoising: A Survey
- Learning a Discriminative Filter Bank within a CNN for Fine-grained Recognition
- Deep Supervision with Shape Concepts for Occlusion-Aware 3D Object Parsing
- AMTnet: Action-Micro-Tube Regression by End-to-end Trainable Deep Architecture
- Self-Supervised Feature Learning by Learning to Spot Artifacts
- Deep Feature Flow for Video Recognition
- Adaptive Feeding: Achieving Fast and Accurate Detections by Adaptively Combining Object Detectors
- Benanza: Automatic Benchmark Generation to Compute "Lower-bound" Latency and Inform Optimizations of Deep Learning Models on GPUs
- Dense Captioning with Joint Inference and Visual Context
- Spatial Feature Calibration and Temporal Fusion for Effective One-stage Video Instance Segmentation
- Attribute-Graph: A Graph based approach to Image Ranking
- Weakly Supervised Object Localization Using Things and Stuff Transfer
- Spatio-temporal Human Action Localisation and Instance Segmentation in Temporally Untrimmed Videos
- Deep Gaussian Conditional Random Field Network: A Model-based Deep Network for Discriminative Denoising
- NETNet: Neighbor Erasing and Transferring Network for Better Single Shot Object Detection
- Oriented Boxes for Accurate Instance Segmentation
- PCNN: Pattern-based Fine-Grained Regular Pruning towards Optimizing CNN Accelerators
- Deep Regionlets for Object Detection
- ProNet: Learning to Propose Object-specific Boxes for Cascaded Neural Networks
- Learning High-level Prior with Convolutional Neural Networks for Semantic Segmentation
- Do More Dropouts in Pool5 Feature Maps for Better Object Detection
- Deep Roto-Translation Scattering for Object Classification
- Dense Recurrent Neural Networks for Scene Labeling
- Learning to detect and localize many objects from few examples
- Learning Detection with Diverse Proposals
- Towards High Performance Video Object Detection
- Learning from Noisy Anchors for One-stage Object Detection
- Attend in groups: a weakly-supervised deep learning framework for learning from web data
- Deep Multi-camera People Detection
- Straight to Shapes: Real-time Detection of Encoded Shapes
- A flexible FPGA accelerator for convolutional neural networks
- Trainable Activation Function in Image Classification
- Sketch2code: Generating a website from a paper mockup
- Better with Less: A Data-Active Perspective on Pre-Training Graph Neural Networks
- Exploit Bounding Box Annotations for Multi-label Object Recognition
- A Coarse-to-Fine Model for 3D Pose Estimation and Sub-category Recognition
- Learning to Segment Moving Objects in Videos
- Analysing domain shift factors between videos and images for object detection
- Parsing Occluded People by Flexible Compositions
- SeGAN: Segmenting and Generating the Invisible
- Collaborative Layer-wise Discriminative Learning in Deep Neural Networks
- Towards Human-Machine Cooperation: Self-supervised Sample Mining for Object Detection
- Spot the Difference by Object Detection
- Improved Deep Learning of Object Category using Pose Information
- AutoScaler: Scale-Attention Networks for Visual Correspondence
- A Systematic Comparison of Deep Learning Architectures in an Autonomous Vehicle
- Action Tubelet Detector for Spatio-Temporal Action Localization
- Generic Tubelet Proposals for Action Localization
- Answering Image Riddles using Vision and Reasoning through Probabilistic Soft Logic
- ViP-CNN: Visual Phrase Guided Convolutional Neural Network
- Pooled Motion Features for First-Person Videos
- Frustum VoxNet for 3D object detection from RGB-D or Depth images
- TensorFlow with user friendly Graphical Framework for object detection API
- End-to-End Integration of a Convolutional Network, Deformable Parts Model and Non-Maximum Suppression
- Learning Action Maps of Large Environments via First-Person Vision
- An Orthogonal-SGD based Learning Approach for MIMO Detection under Multiple Channel Models
- Recovering Spatiotemporal Correspondence between Deformable Objects by Exploiting Consistent Foreground Motion in Video
- LOH and behold: Web-scale visual search, recommendation and clustering using Locally Optimized Hashing
- What can we learn about CNNs from a large scale controlled object dataset?
- Determinantal Point Process as an alternative to NMS
- More Reliable AI Solution: Breast Ultrasound Diagnosis Using Multi-AI Combination
- Deep Reflectance Maps
- Exploiting the Value of the Center-dark Channel Prior for Salient Object Detection
- Probabilistic Oriented Object Detection in Automotive Radar
- Performance of object recognition in wearable videos
- See the Difference: Direct Pre-Image Reconstruction and Pose Estimation by Differentiating HOG
- Instance Scale Normalization for image understanding
- Class Subset Selection for Transfer Learning using Submodularity
- Cascaded Sparse Spatial Bins for Efficient and Effective Generic Object Detection
- Convolutional Networks for Object Category and 3D Pose Estimation from 2D Images
- Learning to Cluster Faces on an Affinity Graph
- Objects as context for detecting their semantic parts
- Feature Selection Convolutional Neural Networks for Visual Tracking
- W-Net: Dense Semantic Segmentation of Subcutaneous Tissue in Ultrasound Images by Expanding U-Net to Incorporate Ultrasound RF Waveform Data
- FA-RPN: Floating Region Proposals for Face Detection
- Scene Parsing via Dense Recurrent Neural Networks with Attentional Selection
- FollowMe: Efficient Online Min-Cost Flow Tracking with Bounded Memory and Computation
- Self-supervised pre-training with acoustic configurations for replay spoofing detection
- Robust Optimization for Deep Regression
- Convolutional Channel Features
- Machine Vision in the Context of Robotics: A Systematic Literature Review
- Unsupervised data augmentation for object detection
- A multilayer backpropagation saliency detection algorithm and its applications
- Augmentation Inside the Network
- Feature Selective Networks for Object Detection
- Articulated motion discovery using pairs of trajectories
- Query-free Clothing Retrieval via Implicit Relevance Feedback
- Better Exploiting OS-CNNs for Better Event Recognition in Images
- Corpus Conversion Service: A machine learning platform to ingest documents at scale [Poster abstract]
- Multi-scale recognition with DAG-CNNs
- Beyond Planar Symmetry: Modeling human perception of reflection and rotation symmetries in the wild
- Semantic Image Cropping
- DeepKey: Towards End-to-End Physical Key Replication From a Single Photograph
- Bounding Box Embedding for Single Shot Person Instance Segmentation
- Cross-domain Image Retrieval with a Dual Attribute-aware Ranking Network
- Subset Feature Learning for Fine-Grained Category Classification
- SHOE: Supervised Hashing with Output Embeddings
- Towards an "In-the-Wild" Emotion Dataset Using a Game-based Framework
- Architecture-aware Network Pruning for Vision Quality Applications
- Non-local RoIs for Instance Segmentation
- Recurrent Residual Module for Fast Inference in Videos
- Triplet-based Deep Similarity Learning for Person Re-Identification
- Learning to Predict the 3D Layout of a Scene
- Using Cross-Model EgoSupervision to Learn Cooperative Basketball Intention
- ROAM: a Rich Object Appearance Model with Application to Rotoscoping
- Compression Artifacts Reduction by a Deep Convolutional Network
- RSAC: Regularized Subspace Approximation Classifier for Lightweight Continuous Learning
- Benchmarking KAZE and MCM for Multiclass Classification
- Active Object Localization with Deep Reinforcement Learning