Fundamental limits to learning closed-form mathematical models from data
arXiv:2204.02704 · doi:10.1038/s41467-023-36657-z
Abstract
Given a finite and noisy dataset generated with a closed-form mathematical model, when is it possible to learn the true generating model from the data alone? This is the question we investigate here. We show that this model-learning problem displays a transition from a low-noise phase in which the true model can be learned, to a phase in which the observation noise is too high for the true model to be learned by any method. Both in the low-noise phase and in the high-noise phase, probabilistic model selection leads to optimal generalization to unseen data. This is in contrast to standard machine learning approaches, including artificial neural networks, which in this particular problem are limited, in the low-noise phase, by their ability to interpolate. In the transition region between the learnable and unlearnable phases, generalization is hard for all approaches including probabilistic model selection.
References in corpus (4)
- Missing and spurious interactions and the reconstruction of complex networks
- Phase transition in the detection of modules in sparse networks
- A Bayesian machine scientist to aid in the solution of challenging scientific problems
- Bayesian machine scientist to compare data collapses for the Nikuradse dataset
Cited by in corpus (5)
- Human mobility is well described by closed-form gravity-like models learned automatically from data
- Transportability without positivity: a synthesis of statistical and simulation modeling
- A Triumvirate of AI Driven Theoretical Discovery
- Using machine learning to find exact analytic solutions to analytically posed physics problems
- The Limits of Inference in Complex Systems: When Stochastic Models Become Indistinguishable