Fixing the train-test resolution discrepancy
arXiv:1906.06423
Abstract
Data-augmentation is key to the training of neural networks for image classification. This paper first shows that existing augmentations induce a significant discrepancy between the typical size of the objects seen by the classifier at train and test time. We experimentally validate that, for a target test resolution, using a lower train resolution offers better classification at test time. We then propose a simple yet effective and efficient strategy to optimize the classifier performance when the train and test resolutions differ. It involves only a computationally cheap fine-tuning of the network at the test resolution. This enables training strong classifiers using small training images. For instance, we obtain 77.1% top-1 accuracy on ImageNet with a ResNet-50 trained on 128x128 images, and 79.8% with one trained on 224x224 image. In addition, if we use extra training data we get 82.5% with the ResNet-50 train with 224x224 images. Conversely, when training a ResNeXt-101 32x48d pre-trained in weakly-supervised fashion on 940 million public images at resolution 224x224 and further optimizing for test resolution 320x320, we obtain a test top-1 accuracy of 86.4% (top-5: 98.0%) (single-crop). To the best of our knowledge this is the highest ImageNet single-crop, top-1 and top-5 accuracy to date.
References in corpus (13)
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
- mixup: Beyond Empirical Risk Minimization
- AutoAugment: Learning Augmentation Policies from Data
- CutMix: Regularization Strategy to Train Strong Classifiers with Localizable Features
- Billion-scale semi-supervised learning for image classification
- Identity Mappings in Deep Residual Networks
- Bag of Tricks for Image Classification with Convolutional Neural Networks
- Fixing the train-test resolution discrepancy: FixEfficientNet
- MultiGrain: a unified image embedding for classes and instances
- Is Second-order Information Helpful for Large-scale Visual Recognition?
- The Herbarium Challenge 2019 Dataset
- The iNaturalist Species Classification and Detection Dataset
Cited by in corpus (11)
- EfficientNetV2: Smaller Models and Faster Training
- Self-training with Noisy Student improves ImageNet classification
- Refiner: Refining Self-attention for Vision Transformers
- MaxUp: A Simple Way to Improve Generalization of Neural Network Training
- Outside the Box: Abstraction-Based Monitoring of Neural Networks
- Making EfficientNet More Efficient: Exploring Batch-Independent Normalization, Group Convolutions and Reduced Resolution Training
- Circumventing Outliers of AutoAugment with Knowledge Distillation
- Ensembles of Spiking Neural Networks
- What it Thinks is Important is Important: Robustness Transfers through Input Gradients
- Temporally Resolution Decrement: Utilizing the Shape Consistency for Higher Computational Efficiency
- Scale Calibrated Training: Improving Generalization of Deep Networks via Scale-Specific Normalization