A Selective Review on Statistical Methods for Massive Data Computation: Distributed Computing, Subsampling, and Minibatch Techniques
arXiv:2403.11163 · doi:10.1080/24754269.2024.2343151
Abstract
This paper presents a selective review of statistical computation methods for massive data analysis. A huge amount of statistical methods for massive data computation have been rapidly developed in the past decades. In this work, we focus on three categories of statistical computation methods: (1) distributed computing, (2) subsampling methods, and (3) minibatch gradient techniques. The first class of literature is about distributed computing and focuses on the situation, where the dataset size is too huge to be comfortably handled by one single computer. In this case, a distributed computation system with multiple computers has to be utilized. The second class of literature is about subsampling methods and concerns about the situation, where the sample size of dataset is small enough to be placed on one single computer but too large to be easily processed by its memory as a whole. The last class of literature studies those minibatch gradient related optimization techniques, which have been extensively used for optimizing various deep learning models.
References in corpus (16)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- One-step sparse estimates in nonconcave penalized likelihood models
- A Stochastic Gradient Method with an Exponential Convergence Rate for Finite Training Sets
- Variants of RMSProp and Adagrad with Logarithmic Regret Bounds
- Understanding the Role of Momentum in Stochastic Gradient Methods
- An Improved Analysis of Stochastic Gradient Descent with Momentum
- Distributed Feature Screening via Componentwise Debiasing
- Divide-and-Conquer Information-Based Optimal Subdata Selection Algorithm
- On Linear Stochastic Approximation: Fine-grained Polyak-Ruppert and Non-Asymptotic Concentration
- An Analysis of Constant Step Size SGD in the Non-convex Regime: Asymptotic Normality and Bias
- DAve-QN: A Distributed Averaged Quasi-Newton Method with Local Superlinear Convergence Rate
- Fast and Robust Sparsity Learning over Networks: A Decentralized Surrogate Median Regression Approach
- Logistic Regression for Massive Data with Rare Events
- Doubly Distributed Supervised Learning and Inference with High-Dimensional Correlated Outcomes
- Decentralised Learning with Random Features and Distributed Gradient Descent