On Statistical Bias In Active Learning: How and When To Fix It
arXiv:2101.11665
Abstract
Active learning is a powerful tool when labelling data is expensive, but it introduces a bias because the training data no longer follows the population distribution. We formalize this bias and investigate the situations in which it can be harmful and sometimes even helpful. We further introduce novel corrective weights to remove bias when doing so is beneficial. Through this, our work not only provides a useful mechanism that can improve the active learning approach, but also an explanation of the empirical successes of various existing approaches which ignore this bias. In particular, we show that this bias can be actively helpful when training overparameterized models -- like neural networks -- with relatively little data.
Published at ICLR 2021 (Spotlight)
References in corpus (12)
- On the Convergence of Adam and Beyond
- Cost-Effective Active Learning for Deep Image Classification
- Bayesian Active Learning for Classification and Preference Learning
- BatchBALD: Efficient and Diverse Batch Acquisition for Deep Bayesian Active Learning
- Bayesian semi-supervised learning for uncertainty-calibrated prediction of molecular properties and active learning
- Half a Percent of Labels is Enough: Efficient Animal Detection in UAV Imagery using Deep CNNs and Active Learning
- Selection via Proxy: Efficient Data Selection for Deep Learning
- Active Learning with Partial Feedback
- Active Learning for Cost-Sensitive Classification
- UPAL: Unbiased Pool Based Active Learning
- Active Learning with Logged Data
- Active Learning for Decision-Making from Imbalanced Observational Data
Cited by in corpus (6)
- A Survey on Active Learning and Human-in-the-Loop Deep Learning for Medical Image Analysis
- Active Testing: Sample-Efficient Model Evaluation
- Causal-BALD: Deep Bayesian Active Learning of Outcomes to Infer Treatment-Effects from Observational Data
- Test Distribution-Aware Active Learning: A Principled Approach Against Distribution Shift and Outliers
- Downstream-Pretext Domain Knowledge Traceback for Active Learning
- ALE: A Simulation-Based Active Learning Evaluation Framework for the Parameter-Driven Comparison of Query Strategies for NLP