A Reliable Effective Terascale Linear Learning System
arXiv:1110.4198
Abstract
We present a system and a set of techniques for learning linear predictors with convex losses on terascale datasets, with trillions of features, {The number of features here refers to the number of non-zero entries in the data matrix.} billions of training examples and millions of parameters in an hour using a cluster of 1000 machines. Individually none of the component techniques are new, but the careful synthesis required to obtain an efficient implementation is. The result is, up to our knowledge, the most scalable and efficient linear learning system reported in the literature (as of 2011 when our experiments were conducted). We describe and thoroughly evaluate the components of the system, showing the importance of the various design choices.
References in corpus (5)
Cited by in corpus (67)
- A Survey on Distributed Machine Learning
- Compressing Neural Networks with the Hashing Trick
- Communication Efficient Distributed Optimization using an Approximate Newton-type Method
- A Reliable Effective Terascale Linear Learning System
- A Linearly-Convergent Stochastic L-BFGS Algorithm
- Decentralized Federated Learning: A Segmented Gossip Approach
- Communication Complexity of Distributed Convex Learning and Optimization
- Making Contextual Decisions with Low Technical Debt
- Fundamental Limits of Online and Distributed Algorithms for Statistical Learning and Estimation
- Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads
- In Defense of MinHash Over SimHash
- Accelerated Mini-Batch Stochastic Dual Coordinate Ascent
- Towards Geo-Distributed Machine Learning
- Scalable Deep Learning on Distributed Infrastructures: Challenges, Techniques and Tools
- Iterative MapReduce for Large Scale Machine Learning
- A Multi-Batch L-BFGS Method for Machine Learning
- Minimax Estimation of Conditional Moment Models
- TuPAQ: An Efficient Planner for Large-scale Predictive Analytic Queries
- Without-Replacement Sampling for Stochastic Gradient Methods: Convergence Results and Application to Distributed Optimization
- A distributed block coordinate descent method for training regularized linear classifiers
- Distributed linear regression by averaging
- Theory of Dual-sparse Regularized Randomized Reduction
- MISSION: Ultra Large-Scale Feature Selection using Count-Sketches
- Probabilistic Graphical Models on Multi-Core CPUs using Java 8
- Distributed Optimization of Multi-Class SVMs
- A Parallel SGD method with Strong Convergence
- A Credit Assignment Compiler for Joint Prediction
- Learning Frames from Text with an Unsupervised Latent Variable Model
- Advances in Asynchronous Parallel and Distributed Optimization
- Quizbowl: The Case for Incremental Question Answering
- Efficient Online Bootstrapping for Large Scale Learning
- Stochastic Gradient Descent on Highly-Parallel Architectures
- Distributed Stochastic Optimization of the Regularized Risk
- Scalable Nonlinear Learning with Adaptive Polynomial Expansions
- An efficient distributed learning algorithm based on effective local functional approximations
- Distributed Training of Structured SVM
- Efficient Communications in Training Large Scale Neural Networks
- An Adaptive Memory Multi-Batch L-BFGS Algorithm for Neural Network Training
- Speculative Approximations for Terascale Analytics
- Supervised Dimensionality Reduction for Big Data
- Dot-Product Join: An Array-Relation Join Operator for Big Model Analytics
- ASAP: Asynchronous Approximate Data-Parallel Computation
- A Distributed Algorithm for Training Nonlinear Kernel Machines
- Para-active learning
- Parallel Stochastic Gradient Descent with Sound Combiners
- Optimal Gradient Checkpoint Search for Arbitrary Computation Graphs
- Improved Algorithms for Agnostic Pool-based Active Classification
- HyperTune: Dynamic Hyperparameter Tuning For Efficient Distribution of DNN Training Over Heterogeneous Systems
- Large-scale Machine Learning for Metagenomics Sequence Classification
- How Much Restricted Isometry is Needed In Nonconvex Matrix Recovery?
- DaSGD: Squeezing SGD Parallelization Performance in Distributed Training Using Delayed Averaging
- Analyzing statistical and computational tradeoffs of estimation procedures
- Stochastic Optimization under Distributional Drift
- Revisiting Large Scale Distributed Machine Learning
- Diverse Online Feature Selection
- Distributed Function Minimization in Apache Spark
- On Second-order Optimization Methods for Federated Learning
- Data-Efficient Methods for Dialogue Systems
- Tell Me Something New: A New Framework for Asynchronous Parallel Learning
- A stochastic coordinate descent primal-dual algorithm with dynamic stepsize for large-scale composite optimization
- A stochastic coordinate descent splitting primal-dual fixed point algorithm and applications to large-scale composite optimization
- Chromatic Learning for Sparse Datasets
- A polynomial expansion line search for large-scale unconstrained minimization of smooth L2-regularized loss functions, with implementation in Apache Spark
- Fast Online "Next Best Offers" using Deep Learning
- Doubly stochastic large scale kernel learning with the empirical kernel map
- A stochastic coordinate descent inertial primal-dual algorithm for large-scale composite optimization
- Greedy Step Averaging: A parameter-free stochastic optimization method