Learning Deep Features for Discriminative Localization
arXiv:1512.04150
Abstract
In this work, we revisit the global average pooling layer proposed in [13], and shed light on how it explicitly enables the convolutional neural network to have remarkable localization ability despite being trained on image-level labels. While this technique was previously proposed as a means for regularizing training, we find that it actually builds a generic localizable deep representation that can be applied to a variety of tasks. Despite the apparent simplicity of global average pooling, we are able to achieve 37.1% top-5 error for object localization on ILSVRC 2014, which is remarkably close to the 34.2% top-5 error achieved by a fully supervised CNN approach. We demonstrate that our network is able to localize the discriminative image regions on a variety of tasks despite not being trained for them
References in corpus (2)
Cited by in corpus (75)
- Simple Baseline for Visual Question Answering
- A feature agnostic approach for glaucoma detection in OCT volumes
- Deformable ConvNets v2: More Deformable, Better Results
- ThiNet: A Filter Level Pruning Method for Deep Neural Network Compression
- What Do We Understand About Convolutional Networks?
- Weakly Supervised Medical Diagnosis and Localization from Multiple Resolutions
- Towards Interpretable Deep Neural Networks by Leveraging Adversarial Examples
- Parameter-Free Spatial Attention Network for Person Re-Identification
- Adversarial Complementary Learning for Weakly Supervised Object Localization
- Two-Stream Neural Networks for Tampered Face Detection
- Morphological classification of galaxies with deep learning: comparing 3-way and 4-way CNNs
- Weakly Supervised Action Localization by Sparse Temporal Pooling Network
- A Taxonomy and Library for Visualizing Learned Features in Convolutional Neural Networks
- Knowing When to Look: Adaptive Attention via A Visual Sentinel for Image Captioning
- SS-CAM: Smoothed Score-CAM for Sharper Visual Feature Localization
- Large-Scale Image Retrieval with Attentive Deep Local Features
- Time Series Classification from Scratch with Deep Neural Networks: A Strong Baseline
- VideoMix: Rethinking Data Augmentation for Video Classification
- Prospects for Theranostics in Neurosurgical Imaging: Empowering Confocal Laser Endomicroscopy Diagnostics via Deep Learning
- ContextLocNet: Context-Aware Deep Network Models for Weakly Supervised Localization
- Two-Phase Learning for Weakly Supervised Object Localization
- Classifying Exoplanet Candidates with Convolutional Neural Networks: Application to the Next Generation Transit Survey
- Chester: A Web Delivered Locally Computed Chest X-Ray Disease Prediction System
- Weakly-Supervised Semantic Segmentation by Iteratively Mining Common Object Features
- Oriented Response Networks
- IS-CAM: Integrated Score-CAM for axiomatic-based explanations
- It Takes Two to Tango: Towards Theory of AI's Mind
- Robust Tumor Localization with Pyramid Grad-CAM
- Action Machine: Rethinking Action Recognition in Trimmed Videos
- Attention-based Context Aggregation Network for Monocular Depth Estimation
- Meta-DETR: Image-Level Few-Shot Object Detection with Inter-Class Correlation Exploitation
- Teaching Categories to Human Learners with Visual Explanations
- Interpretable Convolutional Neural Networks
- CheXpedition: Investigating Generalization Challenges for Translation of Chest X-Ray Algorithms to the Clinical Setting
- Weakly-supervised Visual Grounding of Phrases with Linguistic Structures
- Cardiac Segmentation on CT Images through Shape-Aware Contour Attentions
- Improving Interpretability of Deep Neural Networks with Semantic Information
- Interactively Transferring CNN Patterns for Part Localization
- xCos: An Explainable Cosine Metric for Face Verification Task
- Visual Relationship Detection using Scene Graphs: A Survey
- Learning a Discriminative Filter Bank within a CNN for Fine-grained Recognition
- ProNet: Learning to Propose Object-specific Boxes for Cascaded Neural Networks
- Mining Object Parts from CNNs via Active Question-Answering
- Interpretation of Deep Temporal Representations by Selective Visualization of Internally Activated Nodes
- Improving Object Detection with Inverted Attention
- Explaining Predictions by Approximating the Local Decision Boundary
- Top-down Visual Saliency Guided by Captions
- From Shallow to Deep Interactions Between Knowledge Representation, Reasoning and Machine Learning (Kay R. Amel group)
- Visual Concept Recognition and Localization via Iterative Introspection
- Understanding Convolutional Networks with APPLE : Automatic Patch Pattern Labeling for Explanation
- Segmentation of Surgical Instruments for Minimally-Invasive Robot-Assisted Procedures Using Generative Deep Neural Networks
- Learning to discover and localize visual objects with open vocabulary
- Recurrent U-net: Deep learning to predict daily summertime ozone in the United States
- What and Where: A Context-based Recommendation System for Object Insertion
- Cube Padding for Weakly-Supervised Saliency Prediction in 360° Videos
- Evaluating the performance of the LIME and Grad-CAM explanation methods on a LEGO multi-label image classification task
- Efficient Image Evidence Analysis of CNN Classification Results
- Viraliency: Pooling Local Virality
- Tournament Based Ranking CNN for the Cataract grading
- When Differential Privacy Meets Interpretability: A Case Study
- Thoracic Disease Identification and Localization using Distance Learning and Region Verification
- Recurrent and Spiking Modeling of Sparse Surgical Kinematics
- Classification of Breast Cancer Lesions in Ultrasound Images by using Attention Layer and loss Ensembles in Deep Convolutional Neural Networks
- Single Sample Feature Importance: An Interpretable Algorithm for Low-Level Feature Analysis
- Dense xUnit Networks
- Fast and interpretable classification of small X-ray diffraction datasets using data augmentation and deep neural networks
- Weakly Supervised Foreground Learning for Weakly Supervised Localization and Detection
- DeepWheat: Estimating Phenotypic Traits from Crop Images with Deep Learning
- Beyond Attributes: Adversarial Erasing Embedding Network for Zero-shot Learning
- Train, Diagnose and Fix: Interpretable Approach for Fine-grained Action Recognition
- Egok360: A 360 Egocentric Kinetic Human Activity Video Dataset
- A Gated Peripheral-Foveal Convolutional Neural Network for Unified Image Aesthetic Prediction
- A Comparative Study on Effects of Original and Pseudo Labels for Weakly Supervised Learning for Car Localization Problem
- Self-Improving Semantic Perception for Indoor Localisation
- Scientific Calculator for Designing Trojan Detectors in Neural Networks