Deep Learning is Singular, and That's Good
arXiv:2010.11560 · doi:10.1109/TNNLS.2022.3167409
Abstract
In singular models, the optimal set of parameters forms an analytic set with singularities and classical statistical inference cannot be applied to such models. This is significant for deep learning as neural networks are singular and thus "dividing" by the determinant of the Hessian or employing the Laplace approximation are not appropriate. Despite its potential for addressing fundamental issues in deep learning, singular learning theory appears to have made little inroads into the developing canon of deep learning theory. Via a mix of theory and experiment, we present an invitation to singular learning theory as a vehicle for understanding deep learning and suggest important future work to make singular learning theory directly applicable to how deep learning is performed in practice.
References in corpus (8)
- Scaling Laws for Neural Language Models
- A Widely Applicable Bayesian Information Criterion
- Deep Learning Scaling is Predictable, Empirically
- Generalization Bounds of Stochastic Gradient Descent for Wide and Deep Neural Networks
- Norm-Based Capacity Control in Neural Networks
- Rethinking Parameter Counting in Deep Models: Effective Dimensionality Revisited
- Fractional Langevin Monte Carlo: Exploring Lévy Driven Stochastic Differential Equations for Markov Chain Monte Carlo
- On the interplay between noise and curvature and its effect on optimization and generalization