Information-Theoretic Generalization Bounds for SGLD via Data-Dependent Estimates
arXiv:1911.02151
Abstract
In this work, we improve upon the stepwise analysis of noisy iterative learning algorithms initiated by Pensia, Jog, and Loh (2018) and recently extended by Bu, Zou, and Veeravalli (2019). Our main contributions are significantly improved mutual information bounds for Stochastic Gradient Langevin Dynamics via data-dependent estimates. Our approach is based on the variational characterization of mutual information and the use of data-dependent priors that forecast the mini-batch gradient based on a subset of the training samples. Our approach is broadly applicable within the information-theoretic framework of Russo and Zou (2015) and Xu and Raginsky (2017). Our bound can be tied to a measure of flatness of the empirical risk surface. As compared with other bounds that depend on the squared norms of gradients, empirical investigations show that the terms in our bounds are orders of magnitude smaller.
23 pages, 1 figure. To appear in, Advances in Neural Information Processing Systems (33), 2019
Cited by in corpus (9)
- NeurIPS 2020 Competition: Predicting Generalization in Deep Learning
- Recent advances in deep learning theory
- Shape Matters: Understanding the Implicit Bias of the Noise Covariance
- Information-theoretic generalization bounds for black-box learning algorithms
- A Bayesian Perspective on Training Speed and Model Selection
- Information-Theoretic Analysis of Epistemic Uncertainty in Bayesian Meta-learning
- A Probabilistic Representation of DNNs: Bridging Mutual Information and Generalization
- Probabilistic fine-tuning of pruning masks and PAC-Bayes self-bounded learning
- Optimizing Information-theoretical Generalization Bounds via Anisotropic Noise in SGLD