Active Bias: Training More Accurate Neural Networks by Emphasizing High Variance Samples
arXiv:1704.07433
Abstract
Self-paced learning and hard example mining re-weight training instances to improve learning accuracy. This paper presents two improved alternatives based on lightweight estimates of sample uncertainty in stochastic gradient descent (SGD): the variance in predicted probability of the correct class across iterations of mini-batch SGD, and the proximity of the correct class probability to the decision threshold. Extensive experimental results on six datasets show that our methods reliably improve accuracy in various network architectures, including additional gains on top of other popular training techniques, such as residual learning, momentum, ADAM, batch normalization, dropout, and distillation.
camera-ready version for NIPS 2017
References in corpus (6)
- Distilling the Knowledge in a Neural Network
- Natural Language Processing (almost) from Scratch
- No More Pesky Learning Rates
- A detection threshold in the amplitude spectra calculated from Kepler data obtained during K2 mission
- Learning with Confident Examples: Rank Pruning for Robust Classification with Noisy Labels
- Toward Implicit Sample Noise Modeling: Deviation-driven Matrix Factorization
Cited by in corpus (11)
- MetaKG: Meta-learning on Knowledge Graph for Cold-start Recommendation
- Learning from Noisy Labels for Entity-Centric Information Extraction
- Boosting Facial Expression Recognition by A Semi-Supervised Progressive Teacher
- Dirichlet-Based Prediction Calibration for Learning with Noisy Labels
- Active Mini-Batch Sampling using Repulsive Point Processes
- Distributionally Robust Deep Learning using Hardness Weighted Sampling
- Bayesian Statistics Guided Label Refurbishment Mechanism: Mitigating Label Noise in Medical Image Classification
- Self-Training Sampling with Monolingual Data Uncertainty for Neural Machine Translation
- Informative Sample-Aware Proxy for Deep Metric Learning
- Leveraging Local Variation in Data: Sampling and Weighting Schemes for Supervised Deep Learning
- Mitigating Memorization in Sample Selection for Learning with Noisy Labels