Approximating CNNs with Bag-of-local-Features models works surprisingly well on ImageNet
arXiv:1904.00760
Abstract
Deep Neural Networks (DNNs) excel on many complex perceptual tasks but it has proven notoriously difficult to understand how they reach their decisions. We here introduce a high-performance DNN architecture on ImageNet whose decisions are considerably easier to explain. Our model, a simple variant of the ResNet-50 architecture called BagNet, classifies an image based on the occurrences of small local image features without taking into account their spatial ordering. This strategy is closely related to the bag-of-feature (BoF) models popular before the onset of deep learning and reaches a surprisingly high accuracy on ImageNet (87.6% top-5 for 33 x 33 px features and Alexnet performance for 17 x 17 px features). The constraint on local features makes it straight-forward to analyse how exactly each part of the image influences the classification. Furthermore, the BagNets behave similar to state-of-the art deep neural networks such as VGG-16, ResNet-152 or DenseNet-169 in terms of feature sensitivity, error distribution and interactions between image parts. This suggests that the improvements of DNNs over previous bag-of-feature classifiers in the last few years is mostly achieved by better fine-tuning rather than by qualitatively different decision strategies.
Published as a conference paper at the Seventh International Conference on Learning Representations (ICLR 2019)
Cited by in corpus (31)
- Partial success in closing the gap between human and machine vision
- Batch Normalization is a Cause of Adversarial Vulnerability
- On the Binding Problem in Artificial Neural Networks
- Learning Debiased Representation via Disentangled Feature Augmentation
- PatchGuard++: Efficient Provable Attack Detection against Adversarial Patches
- Efficient Certified Defenses Against Patch Attacks on Image Classifiers
- Are Convolutional Neural Networks or Transformers more like human vision?
- Locality and compositionality in zero-shot learning
- This Looks Like That... Does it? Shortcomings of Latent Space Prototype Interpretability in Deep Networks
- Shape or Texture: Understanding Discriminative Features in CNNs
- ScaleCert: Scalable Certified Defense against Adversarial Patches with Sparse Superficial Layers
- A Graph Neural Network Approach for Scalable Wireless Power Control
- Interpretable and Accurate Fine-grained Recognition via Region Grouping
- KerCNNs: biologically inspired lateral connections for classification of corrupted images
- Positional Encoding as Spatial Inductive Bias in GANs
- Qimera: Data-free Quantization with Synthetic Boundary Supporting Samples
- A Pre-defined Sparse Kernel Based Convolution for Deep CNNs
- Predicting Visual Memory Schemas with Variational Autoencoders
- StyleAugment: Learning Texture De-biased Representations by Style Augmentation without Pre-defined Textures
- Learning compact generalizable neural representations supporting perceptual grouping
- Intentonomy: a Dataset and Study towards Human Intent Understanding
- Recurrent Attention Model with Log-Polar Mapping is Robust against Adversarial Attacks
- Fast Concept Mapping: The Emergence of Human Abilities in Artificial Neural Networks when Learning Embodied and Self-Supervised
- Unsupervised Learning of Multi-level Structures for Anomaly Detection
- Describing Textures using Natural Language
- Assessing The Importance Of Colours For CNNs In Object Recognition
- Cyclic orthogonal convolutions for long-range integration of features
- 4-Connected Shift Residual Networks
- SketchTransfer: A Challenging New Task for Exploring Detail-Invariance and the Abstractions Learned by Deep Networks
- Analyzing the Dependency of ConvNets on Spatial Information
- Learning to Represent and Predict Sets with Deep Neural Networks