Identifying and Compensating for Feature Deviation in Imbalanced Deep Learning
arXiv:2001.01385
Abstract
Classifiers trained with class-imbalanced data are known to perform poorly on test data of the "minor" classes, of which we have insufficient training data. In this paper, we investigate learning a ConvNet classifier under such a scenario. We found that a ConvNet significantly over-fits the minor classes, which is quite opposite to traditional machine learning algorithms that often under-fit minor classes. We conducted a series of analysis and discovered the feature deviation phenomenon -- the learned ConvNet generates deviated features between the training and test data of minor classes -- which explains how over-fitting happens. To compensate for the effect of feature deviation which pushes test data toward low decision value regions, we propose to incorporate class-dependent temperatures (CDT) in training a ConvNet. CDT simulates feature deviation in the training phase, forcing the ConvNet to enlarge the decision values for minor-class data so that it can overcome real feature deviation in the test phase. We validate our approach on benchmark datasets and achieve promising performance. We hope that our insights can inspire new ways of thinking in resolving class-imbalanced deep learning.
References in corpus (3)
Cited by in corpus (10)
- Balanced Meta-Softmax for Long-Tailed Visual Recognition
- Deep Long-Tailed Learning: A Survey
- On Model Calibration for Long-Tailed Object Detection and Instance Segmentation
- Inverse Image Frequency for Long-tailed Image Recognition
- Adversarial Robustness under Long-Tailed Distribution
- Towards Calibrated Model for Long-Tailed Visual Recognition from Prior Perspective
- Alleviating the Incompatibility between Cross Entropy Loss and Episode Training for Few-shot Skin Disease Classification
- MosaicOS: A Simple and Effective Use of Object-Centric Images for Long-Tailed Object Detection
- Investigating Group Distributionally Robust Optimization for Deep Imbalanced Learning: A Case Study of Binary Tabular Data Classification
- Procrustean Training for Imbalanced Deep Learning