Recent advances in deep learning theory
arXiv:2012.10931
Abstract
Deep learning is usually described as an experiment-driven field under continuous criticizes of lacking theoretical foundations. This problem has been partially fixed by a large volume of literature which has so far not been well organized. This paper reviews and organizes the recent advances in deep learning theory. The literature is categorized in six groups: (1) complexity and capacity-based approaches for analyzing the generalizability of deep learning; (2) stochastic differential equations and their dynamic systems for modelling stochastic gradient descent and its variants, which characterize the optimization and generalization of deep learning, partially inspired by Bayesian inference; (3) the geometrical structures of the loss landscape that drives the trajectories of the dynamic systems; (4) the roles of over-parameterization of deep neural networks from both positive and negative perspectives; (5) theoretical foundations of several special structures in network architectures; and (6) the increasingly intensive concerns in ethics and security and their relationships with generalizability.
References in corpus (16)
- Sequence to Sequence Learning with Neural Networks
- Equality of Opportunity in Supervised Learning
- Theoretically Principled Trade-off between Robustness and Accuracy
- The Loss Surfaces of Multilayer Networks
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- Differential Privacy as a Mutual Information Constraint
- Optimization for deep learning: theory and algorithms
- Nested Variational Compression in Deep Gaussian Processes
- Information-Theoretic Generalization Bounds for SGLD via Data-Dependent Estimates
- Lower Bounds on Adversarial Robustness from Optimal Transport
- On Generalization Bounds of a Family of Recurrent Neural Networks
- Estimating Cosmological Parameters from the Dark Matter Distribution
- Generalization Bounds for Convolutional Neural Networks
- Capacity Bounded Differential Privacy
- Tighter Generalization Bounds for Iterative Differentially Private Learning Algorithms
- Efficient hyperparameter optimization by way of PAC-Bayes bound minimization