Deep Learning Theory Review: An Optimal Control and Dynamical Systems Perspective
arXiv:1908.10920
Abstract
Attempts from different disciplines to provide a fundamental understanding of deep learning have advanced rapidly in recent years, yet a unified framework remains relatively limited. In this article, we provide one possible way to align existing branches of deep learning theory through the lens of dynamical system and optimal control. By viewing deep neural networks as discrete-time nonlinear dynamical systems, we can analyze how information propagates through layers using mean field theory. When optimization algorithms are further recast as controllers, the ultimate goal of training processes can be formulated as an optimal control problem. In addition, we can reveal convergence and generalization properties by studying the stochastic dynamics of optimization algorithms. This viewpoint features a wide range of theoretical study from information bottleneck to statistical physics. It also provides a principled way for hyper-parameter tuning when optimal control theory is introduced. Our framework fits nicely with supervised learning and can be extended to other learning problems, such as Bayesian learning, adversarial training, and specific forms of meta learning, without efforts. The review aims to shed lights on the importance of dynamics and optimal control when developing deep learning theory.
Under Submission
References in corpus (8)
- RL: Fast Reinforcement Learning via Slow Reinforcement Learning
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Recasting Gradient-Based Meta-Learning as Hierarchical Bayes
- Training Neural Networks with Local Error Signals
- First-order Methods Almost Always Avoid Saddle Points
- Stochastic Thermodynamics of Learning
- Dynamical Isometry and a Mean Field Theory of LSTMs and GRUs
- Neural Networks with Cheap Differential Operators
Cited by in corpus (7)
- Continuous-in-Depth Neural Networks
- Towards Robust and Stable Deep Learning Algorithms for Forward Backward Stochastic Differential Equations
- Mean-field Langevin System, Optimal Control and Deep Neural Networks
- Finite-Time Convergence of Continuous-Time Optimization Algorithms via Differential Inclusions
- DynamicVAE: Decoupling Reconstruction Error and Disentangled Representation Learning
- A Differential Game Theoretic Neural Optimizer for Training Residual Networks
- A priori guarantees of finite-time convergence for Deep Neural Networks