PsyPhy: A Psychophysics Driven Evaluation Framework for Visual Recognition
arXiv:1611.06448 · doi:10.1109/TPAMI.2018.2849989
Abstract
By providing substantial amounts of data and standardized evaluation protocols, datasets in computer vision have helped fuel advances across all areas of visual recognition. But even in light of breakthrough results on recent benchmarks, it is still fair to ask if our recognition algorithms are doing as well as we think they are. The vision sciences at large make use of a very different evaluation regime known as Visual Psychophysics to study visual perception. Psychophysics is the quantitative examination of the relationships between controlled stimuli and the behavioral responses they elicit in experimental test subjects. Instead of using summary statistics to gauge performance, psychophysics directs us to construct item-response curves made up of individual stimulus responses to find perceptual thresholds, thus allowing one to identify the exact point at which a subject can no longer reliably recognize the stimulus class. In this article, we introduce a comprehensive evaluation framework for visual recognition models that is underpinned by this methodology. Over millions of procedurally rendered 3D scenes and 2D images, we compare the performance of well-known convolutional neural networks. Our results bring into question recent claims of human-like performance, and provide a path forward for correcting newly surfaced algorithmic deficiencies.
9 pages, 4 figures. Published at IEEE Transactions on Pattern Analysis and Machine Intelligence. For supplemental material see http://bjrichardwebster.com/papers/psyphy/supp
References in corpus (5)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Learning a Probabilistic Latent Space of Object Shapes via 3D Generative-Adversarial Modeling
- Comparing deep neural networks against humans: object recognition when the signal gets weaker
- How Deep is the Feature Analysis underlying Rapid Visual Categorization?
- The Neural Representation Benchmark and its Evaluation on Brain and Machine
Cited by in corpus (11)
- Strengths and Weaknesses of Deep Learning Models for Face Recognition Against Image Degradations
- EnlightenGAN: Deep Light Enhancement without Paired Supervision
- Examining the Impact of Blur on Recognition by Convolutional Networks
- Bridging the Gap Between Computational Photography and Visual Recognition
- Crowding Reveals Fundamental Differences in Local vs. Global Processing in Humans and Machines
- NOMAD: A Natural, Occluded, Multi-scale Aerial Dataset, for Emergency Response Scenarios
- Using Synthetic Corruptions to Measure Robustness to Natural Distribution Shifts
- Utilizing Network Properties to Detect Erroneous Inputs
- Towards Robust Pattern Recognition: A Review
- Out-of-Distribution Example Detection in Deep Neural Networks using Distance to Modelled Embedding
- Psychophysical Evaluation of Deep Re-Identification Models