Hamiltonian Descent Methods
arXiv:1809.05042
Abstract
We propose a family of optimization methods that achieve linear convergence using first-order gradient information and constant step sizes on a class of convex functions much larger than the smooth and strongly convex ones. This larger class includes functions whose second derivatives may be singular or unbounded at their minima. Our methods are discretizations of conformal Hamiltonian dynamics, which generalize the classical momentum method to model the motion of a particle with non-standard kinetic energy exposed to a dissipative force and the gradient field of the function of interest. They are first-order in the sense that they require only gradient computation. Yet, crucially the kinetic gradient map can be designed to incorporate information about the convex conjugate in a fashion that allows for linear convergence on convex functions that may be non-smooth or non-strongly convex. We study in detail one implicit and two explicit methods. For one explicit method, we provide conditions under which it converges to stationary points of non-convex functions. For all, we provide conditions on the convex function and kinetic energy pair that guarantee linear convergence, and show that these conditions can be satisfied by functions with power growth. In sum, these methods expand the class of convex functions on which linear convergence is possible with first-order computation.
References in corpus (3)
Cited by in corpus (17)
- Dissecting Neural ODEs
- Implicit Regularization and Momentum Algorithms in Nonlinearly Parameterized Adaptive Control and Prediction
- A Nonsmooth Dynamical Systems Perspective on Accelerated Extensions of ADMM
- Fractional Underdamped Langevin Dynamics: Retargeting SGD with Momentum under Heavy-Tailed Gradient Noise
- Fast, Provably convergent IRLS Algorithm for p-norm Linear Regression
- A Control-Theoretic Perspective on Optimal High-Order Optimization
- Efficient MCMC Sampling with Dimension-Free Convergence Rate using ADMM-type Splitting
- Regret Bounds without Lipschitz Continuity: Online Learning with Relative-Lipschitz Losses
- Optimization with Momentum: Dynamical, Control-Theoretic, and Symplectic Perspectives
- Acceleration in First Order Quasi-strongly Convex Optimization by ODE Discretization
- A Continuous-time Perspective for Modeling Acceleration in Riemannian Optimization
- NEO: Non Equilibrium Sampling on the Orbit of a Deterministic Transform
- FastReg: Fast Non-Rigid Registration via Accelerated Optimisation on the Manifold of Diffeomorphisms
- Quickly Finding a Benign Region via Heavy Ball Momentum in Non-Convex Optimization
- Hamiltonian descent for composite objectives
- Complex Momentum for Optimization in Games
- Resource-Aware Discretization of Accelerated Optimization Flows