First-order Newton-type Estimator for Distributed Estimation and Inference
arXiv:1811.11368 · doi:10.1080/01621459.2021.1891925
Abstract
This paper studies distributed estimation and inference for a general statistical problem with a convex loss that could be non-differentiable. For the purpose of efficient computation, we restrict ourselves to stochastic first-order optimization, which enjoys low per-iteration complexity. To motivate the proposed method, we first investigate the theoretical properties of a straightforward Divide-and-Conquer Stochastic Gradient Descent (DC-SGD) approach. Our theory shows that there is a restriction on the number of machines and this restriction becomes more stringent when the dimension is large. To overcome this limitation, this paper proposes a new multi-round distributed estimation procedure that approximates the Newton step only using stochastic subgradient. The key component in our method is the proposal of a computationally efficient estimator of , where is the population Hessian matrix and is any given vector. Instead of estimating (or ) that usually requires the second-order differentiability of the loss, the proposed First-Order Newton-type Estimator (FONE) directly estimates the vector of interest as a whole and is applicable to non-differentiable losses. Our estimator also facilitates the inference for the empirical risk minimizer. It turns out that the key term in the limiting covariance has the form of , which can be estimated by FONE.
60 pages
References in corpus (5)
- Quantile Regression Under Memory Constraint
- Distributed High-dimensional Regression Under a Quantile Loss Function
- Distributed Inference for Linear Support Vector Machine
- Improved Optimization of Finite Sums with Minibatch Stochastic Variance Reduced Proximal Iterations
- Approximate Newton-based statistical inference using only stochastic gradients
Cited by in corpus (11)
- A Selective Review on Statistical Methods for Massive Data Computation: Distributed Computing, Subsampling, and Minibatch Techniques
- Distributed linear regression by averaging
- Federated Gaussian Process: Convergence, Automatic Personalization and Multi-fidelity Modeling
- Variance Reduced Median-of-Means Estimator for Byzantine-Robust Distributed Inference
- WONDER: Weighted one-shot distributed ridge regression in high dimensions
- Distributed Estimation for Principal Component Analysis: an Enlarged Eigenspace Analysis
- Distributed Estimation and Inference for Semi-parametric Binary Response Models
- MANDERA: Malicious Node Detection in Federated Learning via Ranking
- Distributed Pseudo-Likelihood Method for Community Detection in Large-Scale Networks
- Statistical Estimation and Inference via Local SGD in Federated Learning
- Subsampled One-Step Estimation for Fast Statistical Inference