Deep Networks Can Resemble Human Feed-forward Vision in Invariant Object Recognition
arXiv:1508.03929 · doi:10.1038/srep32672
Abstract
Deep convolutional neural networks (DCNNs) have attracted much attention recently, and have shown to be able to recognize thousands of object categories in natural image databases. Their architecture is somewhat similar to that of the human visual system: both use restricted receptive fields, and a hierarchy of layers which progressively extract more and more abstracted features. Yet it is unknown whether DCNNs match human performance at the task of view-invariant object recognition, whether they make similar errors and use similar representations for this task, and whether the answers depend on the magnitude of the viewpoint variations. To investigate these issues, we benchmarked eight state-of-the-art DCNNs, the HMAX model, and a baseline shallow model and compared their results to those of humans with backward masking. Unlike in all previous DCNN studies, we carefully controlled the magnitude of the viewpoint variations to demonstrate that shallow nets can outperform deep nets and humans when variations are weak. When facing larger variations, however, more layers were needed to match human performance and error distributions, and to have representations that are consistent with human behavior. A very deep net with 18 layers even outperformed humans at the highest variation level, using the most human-like representations.
References in corpus (4)
- Deep Learning in Neural Networks: An Overview
- Deep Neural Networks Reveal a Gradient in the Complexity of Neural Representations across the Brain's Ventral Visual Pathway
- Deep Neural Networks Rival the Representation of Primate IT Cortex for Core Visual Object Recognition
- Bio-inspired Unsupervised Learning of Visual Features Leads to Robust Invariant Object Recognition
Cited by in corpus (17)
- Deep Learning in Spiking Neural Networks
- STDP-based spiking deep convolutional neural networks for object recognition
- Convolutional Neural Networks as a Model of the Visual System: Past, Present, and Future
- Cognitive Science in the era of Artificial Intelligence: A roadmap for reverse-engineering the infant language-learner
- First-spike based visual categorization using reward-modulated STDP
- Transfer Learning for Automated OCTA Detection of Diabetic Retinopathy
- Review: Deep Learning in Electron Microscopy
- Going in circles is the way forward: the role of recurrence in visual inference
- Humans and deep networks largely agree on which kinds of variation make object recognition harder
- Distinguishing mirror from glass: A 'big data' approach to material perception
- Quantifying Legibility of Indoor Spaces Using Deep Convolutional Neural Networks: Case Studies in Train Stations
- If a Human Can See It, So Should Your System: Reliability Requirements for Machine Vision Components
- Modeling biological face recognition with deep convolutional neural networks
- Unveiling the Human-like Similarities of Automatic Facial Expression Recognition: An Empirical Exploration through Explainable AI
- Guiding Visual Attention in Deep Convolutional Neural Networks Based on Human Eye Movements
- Investigating Fine- and Coarse-grained Structural Correspondences Between Deep Neural Networks and Human Object Image Similarity Judgments Using Unsupervised Alignment
- Beyond still images: Temporal features and input variance resilience