Maximizing the information learned from finite data selects a simple model
arXiv:1705.01166 · doi:10.1073/pnas.1715306115
Abstract
We use the language of uninformative Bayesian prior choice to study the selection of appropriately simple effective models. We advocate for the prior which maximizes the mutual information between parameters and predictions, learning as much as possible from limited data. When many parameters are poorly constrained by the available data, we find that this prior puts weight only on boundaries of the parameter manifold. Thus it selects a lower-dimensional effective theory in a principled way, ignoring irrelevant parameter directions. In the limit where there is sufficient data to tightly constrain any number of parameters, this reduces to Jeffreys prior. But we argue that this limit is pathological when applied to the hyper-ribbon parameter manifolds generic in science, because it leads to dramatic dependence on effects invisible to experiment.
9 pages, 8 figures. v3 has improved discussion and adds an appendix about MDL and Bayes factors, and matches version to appear in PNAS (modulo comma placement). Title changed from "Rational Ignorance: Simpler Models Learn More Information from Finite Data"
References in corpus (6)
- Asymptotic Equivalence of Bayes Cross Validation and Widely Applicable Information Criterion in Singular Learning Theory
- A Widely Applicable Bayesian Information Criterion
- Optimal decoding of information from a genetic network
- The sloppy model universality class and the Vandermonde matrix
- Delineating Parameter Unidentifiabilities in Complex Models
- Sloppy nuclear energy density functionals: effective model reduction
Cited by in corpus (13)
- A high-bias, low-variance introduction to Machine Learning for physicists
- Statistical aspects of nuclear mass models
- Causal Geometry
- A scaling law from discrete to continuous solutions of channel capacity problems in the low-noise limit
- Variational Predictive Information Bottleneck
- Unwinding the model manifold: choosing similarity measures to remove local minima in sloppy dynamical systems
- Far from Asymptopia
- Weak transcription factor clustering at binding sites can facilitate information transfer from molecular signals
- Simmering: Sufficient is better than optimal for training neural networks
- Intrinsic regularization effect in Bayesian nonlinear regression scaled by observed data
- Optimizing information transmission in optogenetic Wnt signaling
- Parameters inference and model reduction for the Single-Particle Model of Li ion cells
- On the complexity of logistic regression models