Maximum Principle Based Algorithms for Deep Learning
arXiv:1710.09513
Abstract
The continuous dynamical system approach to deep learning is explored in order to devise alternative frameworks for training algorithms. Training is recast as a control problem and this allows us to formulate necessary optimality conditions in continuous time using the Pontryagin's maximum principle (PMP). A modification of the method of successive approximations is then used to solve the PMP, giving rise to an alternative training algorithm for deep learning. This approach has the advantage that rigorous error estimates and convergence results can be established. We also show that it may avoid some pitfalls of gradient-based methods, such as slow convergence on flat landscapes near saddle points. Furthermore, we demonstrate that it obtains favorable initial convergence rate per-iteration, provided Hamiltonian maximization can be efficiently carried out - a step which is still in need of improvement. Overall, the approach opens up new avenues to attack problems associated with deep learning, such as trapping in slow manifolds and inapplicability of gradient-based methods for discrete trainable variables.
Published version
Cited by in corpus (10)
- A Machine Learning Framework for Solving High-Dimensional Mean Field Game and Mean Field Control Problems
- Algorithms for Solving High Dimensional PDEs: From Nonlinear Monte Carlo to Machine Learning
- End-to-End Quantum Machine Learning Implemented with Controlled Quantum Dynamics
- Deep Learning Approximation of Diffeomorphisms via Linear-Control Systems
- An Optimal Control Method to Compute the Most Likely Transition Path for Stochastic Dynamical Systems with Jumps
- Pontryagin's Minimum Principle and Forward-Backward Sweep Method for the System of HJB-FP Equations in Memory-Limited Partially Observable Stochastic Control
- From NeurODEs to AutoencODEs: a mean-field control framework for width-varying Neural Networks
- A Flow Model of Neural Networks
- A Shooting Formulation of Deep Learning
- Contractivity of the Method of Successive Approximations for Optimal Control