Understanding Deep Image Representations by Inverting Them
arXiv:1412.0035
Abstract
Image representations, from SIFT and Bag of Visual Words to Convolutional Neural Networks (CNNs), are a crucial component of almost any image understanding system. Nevertheless, our understanding of them remains limited. In this paper we conduct a direct analysis of the visual information contained in representations by asking the following question: given an encoding of an image, to which extent is it possible to reconstruct the image itself? To answer this question we contribute a general framework to invert representations. We show that this method can invert representations such as HOG and SIFT more accurately than recent alternatives while being applicable to CNNs too. We then use this technique to study the inverse of recent state-of-the-art CNN image representations for the first time. Among our findings, we show that several layers in CNNs retain photographically accurate information about the image, with different degrees of geometric and photometric invariance.
Cited by in corpus (12)
- Machine Learning of Explicit Order Parameters: From the Ising Model to SU(2) Lattice Gauge Theory
- Amulet: Aggregating Multi-level Convolutional Features for Salient Object Detection
- A deep architecture for unified aesthetic prediction
- Multi-scale Deep Learning Architectures for Person Re-identification
- Interactively Transferring CNN Patterns for Part Localization
- Cooperative Learning with Visual Attributes
- Deep Multi-Modal Image Correspondence Learning
- Context Augmentation for Convolutional Neural Networks
- Train, Diagnose and Fix: Interpretable Approach for Fine-grained Action Recognition
- Constraint-free Natural Image Reconstruction from fMRI Signals Based on Convolutional Neural Network
- Collaborative creativity with Monte-Carlo Tree Search and Convolutional Neural Networks
- Analyzing Learned Convnet Features with Dirichlet Process Gaussian Mixture Models