The importance of better models in stochastic optimization
arXiv:1903.08619 · doi:10.1073/pnas.1908018116
Abstract
Standard stochastic optimization methods are brittle, sensitive to stepsize choices and other algorithmic parameters, and they exhibit instability outside of well-behaved families of objectives. To address these challenges, we investigate models for stochastic minimization and learning problems that exhibit better robustness to problem families and algorithmic parameters. With appropriately accurate models---which we call the aProx family---stochastic methods can be made stable, provably convergent and asymptotically optimal; even modeling that the objective is nonnegative is sufficient for this stability. We extend these results beyond convexity to weakly convex objectives, which include compositions of convex losses with smooth functions common in modern machine learning applications. We highlight the importance of robustness and accurate modeling with a careful experimental evaluation of convergence time and algorithm sensitivity.
References in corpus (5)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Neural Architecture Search with Reinforcement Learning
- Stochastic (Approximate) Proximal Point Methods: Convergence, Optimality, and Adaptivity
- The proximal point method revisited
- Nonasymptotic convergence of stochastic proximal point algorithms for constrained convex optimization
Cited by in corpus (13)
- Protection Against Reconstruction and Its Applications in Private Federated Learning
- Stochastic Polyak Step-size for SGD: An Adaptive Learning Rate for Fast Convergence
- Sensitivity analysis of Wasserstein distributionally robust optimization problems
- Latency considerations for stochastic optimizers in variational quantum algorithms
- Stability and Convergence of Stochastic Gradient Clipping: Beyond Lipschitz Continuity and Smoothness
- Optimizer Benchmarking Needs to Account for Hyperparameter Tuning
- Optimization and Supervised Machine Learning Methods for Fitting Numerical Physics Models without Derivatives
- From low probability to high confidence in stochastic convex optimization
- Stochastic Variance-Reduced Prox-Linear Algorithms for Nonconvex Composite Optimization
- Stochastic optimization over proximally smooth sets
- A Semismooth Newton Stochastic Proximal Point Algorithm with Variance Reduction
- How much progress have we made in neural network training? A New Evaluation Protocol for Benchmarking Optimizers
- Distributed Computation of Stochastic GNE with Partial Information: An Augmented Best-Response Approach