How much does your data exploration overfit? Controlling bias via information usage
arXiv:1511.05219
Abstract
Modern data is messy and high-dimensional, and it is often not clear a priori what are the right questions to ask. Instead, the analyst typically needs to use the data to search for interesting analyses to perform and hypotheses to test. This is an adaptive process, where the choice of analysis to be performed next depends on the results of the previous analyses on the same data. Ultimately, which results are reported can be heavily influenced by the data. It is widely recognized that this process, even if well-intentioned, can lead to biases and false discoveries, contributing to the crisis of reproducibility in science. But while %the adaptive nature of exploration any data-exploration renders standard statistical theory invalid, experience suggests that different types of exploratory analysis can lead to disparate levels of bias, and the degree of bias also depends on the particulars of the data set. In this paper, we propose a general information usage framework to quantify and provably bound the bias and other error metrics of an arbitrary exploratory analysis. We prove that our mutual information based bound is tight in natural settings, and then use it to give rigorous insights into when commonly used procedures do or do not lead to substantially biased estimation. Through the lens of information usage, we analyze the bias of specific exploration procedures such as filtering, rank selection and clustering. Our general framework also naturally motivates randomization techniques that provably reduces exploration bias while preserving the utility of the data analysis. We discuss the connections between our approach and related ideas from differential privacy and blinded data analysis, and supplement our results with illustrative simulations.
Accepted at IEEE Transactions on Information Theory
References in corpus (2)
Cited by in corpus (16)
- Sharpened Generalization Bounds based on Conditional Mutual Information and an Application to Noisy, Iterative Algorithms
- Information Bottleneck and its Applications in Deep Learning
- A Minimax Theory for Adaptive Data Analysis
- Capacity Bounded Differential Privacy
- The Role of Information Complexity and Randomization in Representation Learning
- Challenges in Bayesian Adaptive Data Analysis
- Rate-Distortion Analysis of Minimum Excess Risk in Bayesian Learning
- A Probabilistic Representation of DNNs: Bridging Mutual Information and Generalization
- Do Compressed Representations Generalize Better?
- PAC-Bayesian Transportation Bound
- VizRec: A framework for secure data exploration via visual representation
- Theoretical Guarantees for Model Auditing with Finite Adversaries
- Parameterized Indexed Value Function for Efficient Exploration in Reinforcement Learning
- A Look at the Effect of Sample Design on Generalization through the Lens of Spectral Analysis
- Generalizations of Maximal Inequalities to Arbitrary Selection Rules
- Dependence Measures Bounding the Exploration Bias for General Measurements