Are we done with ImageNet?
arXiv:2006.07159
Abstract
Yes, and no. We ask whether recent progress on the ImageNet classification benchmark continues to represent meaningful generalization, or whether the community has started to overfit to the idiosyncrasies of its labeling procedure. We therefore develop a significantly more robust procedure for collecting human annotations of the ImageNet validation set. Using these new labels, we reassess the accuracy of recently proposed ImageNet classifiers, and find their gains to be substantially smaller than those reported on the original labels. Furthermore, we find the original ImageNet labels to no longer be the best predictors of this independently-collected set, indicating that their usefulness in evaluating vision models may be nearing an end. Nevertheless, we find our annotation procedure to have largely remedied the errors in the original labels, reinforcing ImageNet as a powerful benchmark for future research in visual recognition.
All five authors contributed equally. New labels at https://github.com/google-research/reassessed-imagenet
References in corpus (4)
Cited by in corpus (20)
- MLP-Mixer: An all-MLP Architecture for Vision
- CvT: Introducing Convolutions to Vision Transformers
- Large image datasets: A pyrrhic win for computer vision?
- Refiner: Refining Self-attention for Vision Transformers
- ResNet strikes back: An improved training procedure in timm
- VOLO: Vision Outlooker for Visual Recognition
- Reduced, Reused and Recycled: The Life of a Dataset in Machine Learning Research
- Global Filter Networks for Image Classification
- Self-Damaging Contrastive Learning
- Backward-Compatible Prediction Updates: A Probabilistic Approach
- The Benchmark Lottery
- When does loss-based prioritization fail?
- Consistency Regularization Can Improve Robustness to Label Noise
- A Fast Knowledge Distillation Framework for Visual Recognition
- Divergence Optimization for Noisy Universal Domain Adaptation
- A Tale Of Two Long Tails
- Life is not black and white -- Combining Semi-Supervised Learning with fuzzy labels
- Perspective: Purposeful Failure in Artificial Life and Artificial Intelligence
- Disentangled Variational Information Bottleneck for Multiview Representation Learning
- Exploring and Improving Mobile Level Vision Transformers