Object Detectors Emerge in Deep Scene CNNs
arXiv:1412.6856
Abstract
With the success of new computational architectures for visual processing, such as convolutional neural networks (CNN) and access to image databases with millions of labeled examples (e.g., ImageNet, Places), the state of the art in computer vision is advancing rapidly. One important factor for continued progress is to understand the representations that are learned by the inner layers of these deep architectures. Here we show that object detectors emerge from training CNNs to perform scene classification. As scenes are composed of objects, the CNN for scene classification automatically discovers meaningful objects detectors, representative of the learned scene categories. With object detectors emerging as a result of learning to recognize scenes, our work demonstrates that the same network can perform both scene recognition and object localization in a single forward-pass, without ever having been explicitly taught the notion of objects.
12 pages, ICLR 2015 conference paper
Cited by in corpus (35)
- What makes ImageNet good for transfer learning?
- Do Convolutional Neural Networks Learn Class Hierarchy?
- Look Wider to Match Image Patches with Convolutional Neural Networks
- Direct-Manipulation Visualization of Deep Networks
- See, Hear, and Read: Deep Aligned Representations
- Range Loss for Deep Face Recognition with Long-tail
- Conceptual Spaces for Cognitive Architectures: A Lingua Franca for Different Levels of Representation
- Incorporating Network Built-in Priors in Weakly-supervised Semantic Segmentation
- Visual pathways from the perspective of cost functions and multi-task deep neural networks
- Built-in Foreground/Background Prior for Weakly-Supervised Semantic Segmentation
- Soft Proposal Networks for Weakly Supervised Object Localization
- Deep Contextual Attention for Human-Object Interaction Detection
- Learning to count with deep object features
- Quantifying Legibility of Indoor Spaces Using Deep Convolutional Neural Networks: Case Studies in Train Stations
- Structure-Aware Network for Lane Marker Extraction with Dynamic Vision Sensor
- Funnel Activation for Visual Recognition
- Why my photos look sideways or upside down? Detecting Canonical Orientation of Images using Convolutional Neural Networks
- Multi-Object Classification and Unsupervised Scene Understanding Using Deep Learning Features and Latent Tree Probabilistic Models
- Weakly-supervised localization of diabetic retinopathy lesions in retinal fundus images
- A Study on Multimodal and Interactive Explanations for Visual Question Answering
- Semantics for Global and Local Interpretation of Deep Neural Networks
- ArbiText: Arbitrary-Oriented Text Detection in Unconstrained Scene
- Self-Supervised Visual Place Recognition Learning in Mobile Robots
- Human perception in computer vision
- Deep Multi-Modal Image Correspondence Learning
- Discover and Learn New Objects from Documentaries
- Improved Deep Learning of Object Category using Pose Information
- Deep Visual City Recognition Visualization
- Diverse Sampling for Self-Supervised Learning of Semantic Segmentation
- Mining Interpretable AOG Representations from Convolutional Networks via Active Question Answering
- Optimising the Input Image to Improve Visual Relationship Detection
- Low-Cost Transfer Learning of Face Tasks
- Capturing Localized Image Artifacts through a CNN-based Hyper-image Representation
- Understanding Convolutional Neural Networks with A Mathematical Model
- Fine-Grained Categorization via CNN-Based Automatic Extraction and Integration of Object-Level and Part-Level Features