Do ImageNet Classifiers Generalize to ImageNet?
arXiv:1902.10811
Abstract
We build new test sets for the CIFAR-10 and ImageNet datasets. Both benchmarks have been the focus of intense research for almost a decade, raising the danger of overfitting to excessively re-used test sets. By closely following the original dataset creation processes, we test to what extent current classification models generalize to new data. We evaluate a broad range of models and find accuracy drops of 3% - 15% on CIFAR-10 and 11% - 14% on ImageNet. However, accuracy gains on the original test sets translate to larger gains on the new test sets. Our results suggest that the accuracy drops are not caused by adaptivity, but by the models' inability to generalize to slightly "harder" images than those found in the original test sets.
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- Improved Regularization of Convolutional Neural Networks with Cutout
- Wild Patterns: Ten Years After the Rise of Adversarial Machine Learning
- AutoAugment: Learning Augmentation Policies from Data
- The Ladder: A Reliable Leaderboard for Machine Learning Competitions
Cited by in corpus (139)
- Learning Transferable Visual Models From Natural Language Supervision
- Learning to Prompt for Vision-Language Models
- On the Opportunities and Risks of Foundation Models
- Shortcut Learning in Deep Neural Networks
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
- DINOv2: Learning Robust Visual Features without Supervision
- Reproducible scaling laws for contrastive language-image learning
- Tent: Fully Test-time Adaptation by Entropy Minimization
- WILDS: A Benchmark of in-the-Wild Distribution Shifts
- Learning Robust Global Representations by Penalizing Local Predictive Power
- Out-of-Distribution Generalization via Risk Extrapolation (REx)
- Self-training with Noisy Student improves ImageNet classification
- EVA-02: A Visual Representation for Neon Genesis
- Ten Years of Generative Adversarial Nets (GANs): A survey of the state-of-the-art
- Improving robustness against common corruptions by covariate shift adaptation
- Image Representations Learned With Unsupervised Pre-Training Contain Human-like Biases
- Swin Transformer V2: Scaling Up Capacity and Resolution
- Accountability in an Algorithmic Society: Relationality, Responsibility, and Robustness in Machine Learning
- Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations
- Evaluating Prediction-Time Batch Normalization for Robustness under Covariate Shift
- Improving Robustness Without Sacrificing Accuracy with Patch Gaussian Augmentation
- LeViT: a Vision Transformer in ConvNet's Clothing for Faster Inference
- The Risks of Invariant Risk Minimization
- A Fourier Perspective on Model Robustness in Computer Vision
- Enhanced Convolutional Neural Tangent Kernels
- Are we done with ImageNet?
- Fixing the train-test resolution discrepancy: FixEfficientNet
- Run-Time Monitoring of Machine Learning for Robotic Perception: A Survey of Emerging Trends
- Cold Case: The Lost MNIST Digits
- Self-Supervised Policy Adaptation during Deployment
- Discovering and Validating AI Errors With Crowdsourced Failure Reports
- Data and its (dis)contents: A survey of dataset development and use in machine learning research
- A Survey of Deep Learning for Scientific Discovery
- Long-Short Transformer: Efficient Transformers for Language and Vision
- How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers
- Frustratingly Simple Domain Generalization via Image Stylization
- Contrastive Training for Improved Out-of-Distribution Detection
- Measuring Robustness in Deep Learning Based Compressive Sensing
- Image Classification with Small Datasets: Overview and Benchmark
- Identifying Mislabeled Data using the Area Under the Margin Ranking
- Large image datasets: A pyrrhic win for computer vision?
- Open-Set Recognition: a Good Closed-Set Classifier is All You Need?
- Real Risks of Fake Data: Synthetic Data, Diversity-Washing and Consent Circumvention
- Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis
- PSLT: A Light-weight Vision Transformer with Ladder Self-Attention and Progressive Shift
- Improving the Transferability of Adversarial Examples with Arbitrary Style Transfer
- Is it enough to optimize CNN architectures on ImageNet?
- Assaying Out-Of-Distribution Generalization in Transfer Learning
- Mitigating Bias in Calibration Error Estimation
- Can Fairness be Automated? Guidelines and Opportunities for Fairness-aware AutoML
- Deep regularization and direct training of the inner layers of Neural Networks with Kernel Flows
- Robustness properties of Facebook's ResNeXt WSL models
- Poisoning the Unlabeled Dataset of Semi-Supervised Learning
- Evaluating Weakly Supervised Object Localization Methods Right
- Angler: Helping Machine Translation Practitioners Prioritize Model Improvements
- Empirical Upper Bound in Object Detection and More
- Poisoning and Backdooring Contrastive Learning
- HARK Side of Deep Learning -- From Grad Student Descent to Automated Machine Learning
- Image Based Identification of Ghanaian Timbers Using the XyloTron: Opportunities, Risks and Challenges
- Multi-Objective Evolutionary Design of Deep Convolutional Neural Networks for Image Classification
- Understanding Isomorphism Bias in Graph Data Sets
- Fourier-basis Functions to Bridge Augmentation Gap: Rethinking Frequency Augmentation in Image Classification
- Designing Ocean Vision AI: An Investigation of Community Needs for Imaging-based Ocean Conservation
- SimMIM: A Simple Framework for Masked Image Modeling
- Multi-Scale and Multi-Layer Contrastive Learning for Domain Generalization
- The Cells Out of Sample (COOS) dataset and benchmarks for measuring out-of-sample generalization of image classifiers
- MT3: Meta Test-Time Training for Self-Supervised Test-Time Adaption
- A Generalizable and Accessible Approach to Machine Learning with Global Satellite Imagery
- Optimizing JPEG Quantization for Classification Networks
- Do Offline Metrics Predict Online Performance in Recommender Systems?
- Subhalo effective density slope measurements from HST strong lensing data with neural likelihood-ratio estimation
- On Interaction Between Augmentations and Corruptions in Natural Corruption Robustness
- Angular Visual Hardness
- Environment Inference for Invariant Learning
- Discrete Representations Strengthen Vision Transformer Robustness
- Deep Ensembles for Low-Data Transfer Learning
- VisDA-2021 Competition Universal Domain Adaptation to Improve Performance on Out-of-Distribution Data
- On Mixup Regularization
- ObjectNet Dataset: Reanalysis and Correction
- On Synthetic Data for Back Translation
- Adaptive Label Smoothing
- Detecting Overfitting via Adversarial Examples
- Domain-Aware Dynamic Networks
- An Automotive Case Study on the Limits of Approximation for Object Detection
- Weakly-supervised Object Localization for Few-shot Learning and Fine-grained Few-shot Learning
- A Bayesian Approach to OOD Robustness in Image Classification
- MLDemon: Deployment Monitoring for Machine Learning Systems
- Robustness in Compressed Neural Networks for Object Detection
- Re-labeling ImageNet: from Single to Multi-Labels, from Global to Localized Labels
- No One Representation to Rule Them All: Overlapping Features of Training Methods
- Learning Implicitly with Noisy Data in Linear Arithmetic
- KinePose: A temporally optimized inverse kinematics technique for 6DOF human pose estimation with biomechanical constraints
- Ensemble Model Patching: A Parameter-Efficient Variational Bayesian Neural Network
- Giving Up Control: Neurons as Reinforcement Learning Agents
- Pyramid Adversarial Training Improves ViT Performance
- Low-Rank Training of Deep Neural Networks for Emerging Memory Technology
- On Calibration of Mixup Training for Deep Neural Networks
- An empirical study of pretrained representations for few-shot classification
- Curriculum By Smoothing
- A Fine-Grained Analysis on Distribution Shift
- NEURO-DRAM: a 3D recurrent visual attention model for interpretable neuroimaging classification
- CNN Feature Map Augmentation for Single-Source Domain Generalization
- Trapped in texture bias? A large scale comparison of deep instance segmentation
- Defending Against Image Corruptions Through Adversarial Augmentations
- Learning Visual Representations for Transfer Learning by Suppressing Texture
- Adversarial Scrutiny of Evidentiary Statistical Software
- Global Multiclass Classification and Dataset Construction via Heterogeneous Local Experts
- Using Synthetic Corruptions to Measure Robustness to Natural Distribution Shifts
- Generalized Resubstitution for Classification Error Estimation
- A Modular System for Enhanced Robustness of Multimedia Understanding Networks via Deep Parametric Estimation
- Variational Resampling Based Assessment of Deep Neural Networks under Distribution Shift
- Compressive Visual Representations
- RATT: Leveraging Unlabeled Data to Guarantee Generalization
- Rip van Winkle's Razor: A Simple Estimate of Overfit to Test Data
- Fortify Machine Learning Production Systems: Detect and Classify Adversarial Attacks
- The Unreasonable Effectiveness of Patches in Deep Convolutional Kernels Methods
- Characterizing Generalization under Out-Of-Distribution Shifts in Deep Metric Learning
- ResIST: Layer-Wise Decomposition of ResNets for Distributed Training
- Trivial or impossible -- dichotomous data difficulty masks model differences (on ImageNet and beyond)
- Why Do Better Loss Functions Lead to Less Transferable Features?
- Geometry matters: Exploring language examples at the decision boundary
- Robust, Extensible, and Fast: Teamed Classifiers for Vehicle Tracking and Vehicle Re-ID in Multi-Camera Networks
- FOCUS: Familiar Objects in Common and Uncommon Settings
- Sparse MoEs meet Efficient Ensembles
- Natural Adversarial Objects
- Active Online Learning with Hidden Shifting Domains
- Meta Two-Sample Testing: Learning Kernels for Testing with Limited Data
- Adaptive Test-Time Augmentation for Low-Power CPU
- Creative Captioning: An AI Grand Challenge Based on the Dixit Board Game
- Optimal multiclass overfitting by sequence reconstruction from Hamming queries
- Out-of-Distribution Generalization in Kernel Regression
- Covariate Shift in High-Dimensional Random Feature Regression
- Reappraising Domain Generalization in Neural Networks
- Establishing an Evaluation Metric to Quantify Climate Change Image Realism
- On Deep Neural Network Calibration by Regularization and its Impact on Refinement
- Generalization over different cellular automata rules learned by a deep feed-forward neural network
- Characterizing and Improving the Robustness of Self-Supervised Learning through Background Augmentations
- How Well Do Sparse Imagenet Models Transfer?