A review of distributed statistical inference
arXiv:2304.06245 · doi:10.1080/24754269.2021.1974158
Abstract
The rapid emergence of massive datasets in various fields poses a serious challenge to traditional statistical methods. Meanwhile, it provides opportunities for researchers to develop novel algorithms. Inspired by the idea of divide-and-conquer, various distributed frameworks for statistical estimation and inference have been proposed. They were developed to deal with large-scale statistical optimization problems. This paper aims to provide a comprehensive review for related literature. It includes parametric models, nonparametric models, and other frequently used models. Their key ideas and theoretical properties are summarized. The trade-off between communication cost and estimate precision together with other concerns are discussed.
References in corpus (8)
- Distributed learning with regularized least squares
- Optimality guarantees for distributed statistical estimation
- Distributed Estimation, Information Loss and Exponential Families
- Distributed Feature Screening via Componentwise Debiasing
- A General Framework for Robust Testing and Confidence Regions in High-Dimensional Quantile Regression
- Scalable and Efficient Statistical Inference with Estimating Functions in the MapReduce Paradigm for Big Data
- Rates of Convergence for Large-scale Nearest Neighbor Classification
- Efficient Estimation for Generalized Linear Models on a Distributed System with Nonrandomly Distributed Data