Object Detectors Emerge in Deep Scene CNNs
arXiv:1412.6856
Abstract
With the success of new computational architectures for visual processing, such as convolutional neural networks (CNN) and access to image databases with millions of labeled examples (e.g., ImageNet, Places), the state of the art in computer vision is advancing rapidly. One important factor for continued progress is to understand the representations that are learned by the inner layers of these deep architectures. Here we show that object detectors emerge from training CNNs to perform scene classification. As scenes are composed of objects, the CNN for scene classification automatically discovers meaningful objects detectors, representative of the learned scene categories. With object detectors emerging as a result of learning to recognize scenes, our work demonstrates that the same network can perform both scene recognition and object localization in a single forward-pass, without ever having been explicitly taught the notion of objects.
12 pages, ICLR 2015 conference paper
Cited by in corpus (54)
- Anabranch Network for Camouflaged Object Segmentation
- What makes ImageNet good for transfer learning?
- Do Convolutional Neural Networks Learn Class Hierarchy?
- Look Wider to Match Image Patches with Convolutional Neural Networks
- Direct-Manipulation Visualization of Deep Networks
- See, Hear, and Read: Deep Aligned Representations
- Range Loss for Deep Face Recognition with Long-tail
- Conceptual Spaces for Cognitive Architectures: A Lingua Franca for Different Levels of Representation
- Incorporating Network Built-in Priors in Weakly-supervised Semantic Segmentation
- Visual pathways from the perspective of cost functions and multi-task deep neural networks
- Built-in Foreground/Background Prior for Weakly-Supervised Semantic Segmentation
- Learning to count with deep object features
- Deep Contextual Attention for Human-Object Interaction Detection
- Soft Proposal Networks for Weakly Supervised Object Localization
- Global Aggregation then Local Distribution for Scene Parsing
- Quantifying Legibility of Indoor Spaces Using Deep Convolutional Neural Networks: Case Studies in Train Stations
- Structure-Aware Network for Lane Marker Extraction with Dynamic Vision Sensor
- Funnel Activation for Visual Recognition
- Why my photos look sideways or upside down? Detecting Canonical Orientation of Images using Convolutional Neural Networks
- Visual Relationship Detection using Scene Graphs: A Survey
- Multi-Object Classification and Unsupervised Scene Understanding Using Deep Learning Features and Latent Tree Probabilistic Models
- Weakly-supervised localization of diabetic retinopathy lesions in retinal fundus images
- A Study on Multimodal and Interactive Explanations for Visual Question Answering
- Semantics for Global and Local Interpretation of Deep Neural Networks
- Self-Supervised Visual Place Recognition Learning in Mobile Robots
- Explainable Abstract Trains Dataset
- ArbiText: Arbitrary-Oriented Text Detection in Unconstrained Scene
- Learning Interpretable Concept Groups in CNNs
- Human perception in computer vision
- Discover and Learn New Objects from Documentaries
- Fast Crack Detection Using Convolutional Neural Network
- Deep Multi-Modal Image Correspondence Learning
- Perspective: A Phase Diagram for Deep Learning unifying Jamming, Feature Learning and Lazy Training
- Quantifying Learnability and Describability of Visual Concepts Emerging in Representation Learning
- Deep Learning for Robust Motion Segmentation with Non-Static Cameras
- Improved Deep Learning of Object Category using Pose Information
- Deep Visual City Recognition Visualization
- A Selective Survey on Versatile Knowledge Distillation Paradigm for Neural Network Models
- Finding Discriminative Filters for Specific Degradations in Blind Super-Resolution
- Diverse Sampling for Self-Supervised Learning of Semantic Segmentation
- Mining Interpretable AOG Representations from Convolutional Networks via Active Question Answering
- On the Evolution of Neuron Communities in a Deep Learning Architecture
- Understanding Character Recognition using Visual Explanations Derived from the Human Visual System and Deep Networks
- Integrated Grad-CAM: Sensitivity-Aware Visual Explanation of Deep Convolutional Networks via Integrated Gradient-Based Scoring
- Optimising the Input Image to Improve Visual Relationship Detection
- Toward a Realistic Benchmark for Out-of-Distribution Detection
- Capturing Localized Image Artifacts through a CNN-based Hyper-image Representation
- Low-Cost Transfer Learning of Face Tasks
- Understanding Convolutional Neural Networks with A Mathematical Model
- AMC-Loss: Angular Margin Contrastive Loss for Improved Explainability in Image Classification
- Interpretable Attention Guided Network for Fine-grained Visual Classification
- Ada-SISE: Adaptive Semantic Input Sampling for Efficient Explanation of Convolutional Neural Networks
- Interpretability in Safety-Critical FinancialTrading Systems
- Fine-Grained Categorization via CNN-Based Automatic Extraction and Integration of Object-Level and Part-Level Features