A note on regularised NTK dynamics with an application to PAC-Bayesian training
arXiv:2312.13259
Abstract
We establish explicit dynamics for neural networks whose training objective has a regularising term that constrains the parameters to remain close to their initial value. This keeps the network in a lazy training regime, where the dynamics can be linearised around the initialisation. The standard neural tangent kernel (NTK) governs the evolution during the training in the infinite-width limit, although the regularisation yields an additional term appears in the differential equation describing the dynamics. This setting provides an appropriate framework to study the evolution of wide networks trained to optimise generalisation objectives such as PAC-Bayes bounds, and hence potentially contribute to a deeper theoretical understanding of such networks.
References in corpus (10)
- A Note on the PAC Bayesian Theorem
- Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
- On the properties of variational approximations of Gibbs posteriors
- Finite Versus Infinite Neural Networks: an Empirical Study
- Feature Learning in Infinite-Width Neural Networks
- Learning PAC-Bayes Priors for Probabilistic Neural Networks
- PAC-Bayes Compression Bounds So Tight That They Can Explain Generalization
- Characterizing the Spectrum of the NTK via a Power Series Expansion
- Learning via Wasserstein-Based High Probability Generalisation Bounds
- Demystify Optimization and Generalization of Over-parameterized PAC-Bayesian Learning