Statistical Inference for the Population Landscape via Moment Adjusted Stochastic Gradients
arXiv:1712.07519 · doi:10.1111/rssb.12313
Abstract
Modern statistical inference tasks often require iterative optimization methods to compute the solution. Convergence analysis from an optimization viewpoint only informs us how well the solution is approximated numerically but overlooks the sampling nature of the data. In contrast, recognizing the randomness in the data, statisticians are keen to provide uncertainty quantification, or confidence, for the solution obtained using iterative optimization methods. This paper makes progress along this direction by introducing the moment-adjusted stochastic gradient descents, a new stochastic optimization method for statistical inference. We establish non-asymptotic theory that characterizes the statistical distribution for certain iterative methods with optimization guarantees. On the statistical front, the theory allows for model mis-specification, with very mild conditions on the data. For optimization, the theory is flexible for both convex and non-convex cases. Remarkably, the moment-adjusting idea motivated from "error standardization" in statistics achieves a similar effect as acceleration in first-order optimization methods used to fit generalized linear models. We also demonstrate this acceleration effect in the non-convex setting through numerical experiments.
Journal of the Royal Statistical Society: Series B (Statistical Methodology) 2019, to appear
References in corpus (6)
- A Stochastic Gradient Method with an Exponential Convergence Rate for Finite Training Sets
- Artificial Intelligence and Statistics
- Non-convex learning via Stochastic Gradient Langevin Dynamics: a nonasymptotic analysis
- Global Convergence of Online Limited Memory BFGS
- The Landscape of Empirical Risk for Non-convex Losses
- Newton Sketch: A Linear-time Optimization Algorithm with Linear-Quadratic Convergence
Cited by in corpus (9)
- HiGrad: Uncertainty Quantification for Online Learning and Stochastic Approximation
- On Linear Stochastic Approximation: Fine-grained Polyak-Ruppert and Non-Asymptotic Concentration
- Fast and Robust Online Inference with Stochastic Gradient Descent via Random Scaling
- Online Covariance Matrix Estimation in Stochastic Gradient Descent
- Inference by Stochastic Optimization: A Free-Lunch Bootstrap
- Online Statistical Inference for Stochastic Optimization via Kiefer-Wolfowitz Methods
- Statistical Inference with Local Optima
- Statistical Estimation and Inference via Local SGD in Federated Learning
- Online Tensor Inference