HOGWILD!: A Lock-Free Approach to Parallelizing Stochastic Gradient Descent
arXiv:1106.5730
Abstract
Stochastic Gradient Descent (SGD) is a popular algorithm that can achieve state-of-the-art performance on a variety of machine learning tasks. Several researchers have recently proposed schemes to parallelize SGD, but all require performance-destroying memory locking and synchronization. This work aims to show using novel theoretical analysis, algorithms, and implementation that SGD can be implemented without any locking. We present an update scheme called HOGWILD! which allows processors access to shared memory with the possibility of overwriting each other's work. We show that when the associated optimization problem is sparse, meaning most gradient updates only modify small parts of the decision variable, then HOGWILD! achieves a nearly optimal rate of convergence. We demonstrate experimentally that HOGWILD! outperforms alternative schemes that use locking by an order of magnitude.
22 pages, 10 figures
References in corpus (1)
Cited by in corpus (426)
- TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems
- TensorFlow: A system for large-scale machine learning
- An overview of gradient descent optimization algorithms
- Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
- Asynchronous Methods for Deep Reinforcement Learning
- Federated Optimization: Distributed Machine Learning for On-Device Intelligence
- Deep Learning with Limited Numerical Precision
- Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training
- Speeding Up Distributed Machine Learning Using Codes
- Revisiting Distributed Synchronous SGD
- Don't Decay the Learning Rate, Increase the Batch Size
- Large-scale Simple Question Answering with Memory Networks
- Group Sparse Regularization for Deep Neural Networks
- Can Decentralized Algorithms Outperform Centralized Algorithms? A Case Study for Decentralized Parallel Stochastic Gradient Descent
- Communication Efficient Distributed Optimization using an Approximate Newton-type Method
- Adaptive Computation Time for Recurrent Neural Networks
- Asynchronous Parallel Stochastic Gradient for Nonconvex Optimization
- DyNet: The Dynamic Neural Network Toolkit
- Recent Advances in Convolutional Neural Networks
- Convex Optimization for Big Data
- Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions
- Analogical Inference for Multi-Relational Embeddings
- Asynchronous Distributed ADMM for Large-Scale Optimization- Part I: Algorithm and Convergence Analysis
- PyTorch-BigGraph: A Large-scale Graph Embedding System
- node2vec: Scalable Feature Learning for Networks
- Federated Collaborative Filtering for Privacy-Preserving Personalized Recommendation System
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- MD-GAN: Multi-Discriminator Generative Adversarial Networks for Distributed Datasets
- Local SGD Converges Fast and Communicates Little
- Gradient Sparsification for Communication-Efficient Distributed Optimization
- A Bi-layered Parallel Training Architecture for Large-scale Convolutional Neural Networks
- Inferring Algorithmic Patterns with Stack-Augmented Recurrent Nets
- A Berkeley View of Systems Challenges for AI
- Poincaré Embeddings for Learning Hierarchical Representations
- Cooperative SGD: A unified Framework for the Design and Analysis of Communication-Efficient SGD Algorithms
- Deep Learning in the Automotive Industry: Applications and Tools
- Asynchronous Stochastic Gradient Descent with Delay Compensation
- Robust Federated Learning in a Heterogeneous Environment
- An Asynchronous Parallel Stochastic Coordinate Descent Algorithm
- Metadata Embeddings for User and Item Cold-start Recommendations
- Deep Learning in Mobile and Wireless Networking: A Survey
- Communication Complexity of Distributed Convex Learning and Optimization
- Parallel training of DNNs with Natural Gradient and Parameter Averaging
- Streaming Variational Bayes
- Question Answering with Subgraph Embeddings
- Neural SLAM: Learning to Explore with External Memory
- Factoring nonnegative matrices with linear programs
- Communication-Efficient Distributed Dual Coordinate Ascent
- Deep Learning of Representations: Looking Forward
- Taming the Wild: A Unified Analysis of Hogwild!-Style Algorithms
- AIDE: Fast and Communication Efficient Distributed Optimization
- GraphVite: A High-Performance CPU-GPU Hybrid System for Node Embedding
- On Variance Reduction in Stochastic Gradient Descent and its Asynchronous Variants
- Why Random Reshuffling Beats Stochastic Gradient Descent
- Accelerating Federated Learning over Reliability-Agnostic Clients in Mobile Edge Computing Systems
- Heterogeneous Information Network Embedding for Meta Path based Proximity
- The Non-IID Data Quagmire of Decentralized Machine Learning
- Variance Reduction in SGD by Distributed Importance Sampling
- Open Question Answering with Weakly Supervised Embedding Models
- Efficient Parallel Methods for Deep Reinforcement Learning
- Coordinate Friendly Structures, Algorithms and Applications
- The Error-Feedback Framework: Better Rates for SGD with Delayed Gradients and Compressed Communication
- Applying Deep Learning to Answer Selection: A Study and An Open Task
- Fundamental Limits of Online and Distributed Algorithms for Statistical Learning and Estimation
- High-Performance Distributed ML at Scale through Parameter Server Consistency Models
- Scaling Deep Learning on GPU and Knights Landing clusters
- GPU Asynchronous Stochastic Gradient Descent to Speed Up Neural Network Training
- Asynchronous Distributed ADMM for Large-Scale Optimization- Part II: Linear Convergence Analysis and Numerical Performance
- Communication optimization strategies for distributed deep neural network training: A survey
- Asynchronous Decentralized Parallel Stochastic Gradient Descent
- Natural Compression for Distributed Deep Learning
- Distributed Representations of Signed Networks
- Deep learning with Elastic Averaging SGD
- Adding vs. Averaging in Distributed Primal-Dual Optimization
- Accelerated Mini-Batch Stochastic Dual Coordinate Ascent
- BoostClean: Automated Error Detection and Repair for Machine Learning
- Parallel Successive Convex Approximation for Nonsmooth Nonconvex Optimization
- mvn2vec: Preservation and Collaboration in Multi-View Network Embedding
- Minimizing Latency for Secure Coded Computing Using Secret Sharing via Staircase Codes
- Swivel: Improving Embeddings by Noticing What's Missing
- Exploring Student Check-In Behavior for Improved Point-of-Interest Prediction
- On Nonconvex Optimization for Machine Learning: Gradients, Stochasticity, and Saddle Points
- Strong error analysis for stochastic gradient descent optimization algorithms
- Stochastic Dual Ascent for Solving Linear Systems
- rlpyt: A Research Code Base for Deep Reinforcement Learning in PyTorch
- Occupy the Cloud: Distributed Computing for the 99%
- Distributed Block Coordinate Descent for Minimizing Partially Separable Functions
- Optimizing Network Performance for Distributed DNN Training on GPU Clusters: ImageNet/AlexNet Training in 1.5 Minutes
- Omnivore: An Optimizer for Multi-device Deep Learning on CPUs and GPUs
- Asynchronous and Parallel Distributed Pose Graph Optimization
- Building a Large-scale Multimodal Knowledge Base System for Answering Visual Queries
- Polynomially Coded Regression: Optimal Straggler Mitigation via Data Encoding
- ImageNet Training in Minutes
- Scalable Deep Learning on Distributed Infrastructures: Challenges, Techniques and Tools
- Addressing Overfitting on Pointcloud Classification using Atrous XCRF
- Effective Diversity in Population Based Reinforcement Learning
- An Asynchronous Parallel Randomized Kaczmarz Algorithm
- Improved asynchronous parallel optimization analysis for stochastic incremental methods
- ErasureHead: Distributed Gradient Descent without Delays Using Approximate Gradient Coding
- A Cost-based Optimizer for Gradient Descent Optimization
- PASSCoDe: Parallel ASynchronous Stochastic dual Co-ordinate Descent
- RedSync : Reducing Synchronization Traffic for Distributed Deep Learning
- Asynchronous stochastic convex optimization
- Federated Learning with Buffered Asynchronous Aggregation
- Faasm: Lightweight Isolation for Efficient Stateful Serverless Computing
- Stochastic Polyak Step-size for SGD: An Adaptive Learning Rate for Fast Convergence
- Feature Engineering for Knowledge Base Construction
- The Implicit Regularization of Stochastic Gradient Flow for Least Squares
- SGD: General Analysis and Improved Rates
- CONDENSE: A Reconfigurable Knowledge Acquisition Architecture for Future 5G IoT
- Asynchronous Gibbs Sampling
- Stochastic, Distributed and Federated Optimization for Machine Learning
- Federated Learning With Quantized Global Model Updates
- Breaking the Communication-Privacy-Accuracy Trilemma
- Nonconvex Sparse Learning via Stochastic Optimization with Progressive Variance Reduction
- Preference Completion: Large-scale Collaborative Ranking from Pairwise Comparisons
- CYCLADES: Conflict-free Asynchronous Machine Learning
- Private Empirical Risk Minimization Beyond the Worst Case: The Effect of the Constraint Set Geometry
- Parallel and Distributed Block-Coordinate Frank-Wolfe Algorithms
- GIANT: Globally Improved Approximate Newton Method for Distributed Optimization
- Analysis and Implementation of an Asynchronous Optimization Algorithm for the Parameter Server
- HiGrad: Uncertainty Quantification for Online Learning and Stochastic Approximation
- SpreadGNN: Serverless Multi-task Federated Learning for Graph Neural Networks
- AMPNet: Asynchronous Model-Parallel Training for Dynamic Neural Networks
- Factorbird - a Parameter Server Approach to Distributed Matrix Factorization
- The Sound of APALM Clapping: Faster Nonsmooth Nonconvex Optimization with Stochastic Asynchronous PALM
- PipeMare: Asynchronous Pipeline Parallel DNN Training
- CHAOS: A Parallelization Scheme for Training Convolutional Neural Networks on Intel Xeon Phi
- A Multi-Batch L-BFGS Method for Machine Learning
- Election Coding for Distributed Learning: Protecting SignSGD against Byzantine Attacks
- BlinkML: Efficient Maximum Likelihood Estimation with Probabilistic Guarantees
- Skip-gram word embeddings in hyperbolic space
- Deep Leakage from Gradients
- Convergence of Distributed Stochastic Variance Reduced Methods without Sampling Extra Data
- Programming with Personalized PageRank: A Locally Groundable First-Order Probabilistic Logic
- The Asynchronous PALM Algorithm for Nonsmooth Nonconvex Problems
- CosmoFlow: Using Deep Learning to Learn the Universe at Scale
- Accelerating Generalized Linear Models with MLWeaving: A One-Size-Fits-All System for Any-precision Learning (Technical Report)
- Caffe con Troll: Shallow Ideas to Speed Up Deep Learning
- DistDGL: Distributed Graph Neural Network Training for Billion-Scale Graphs
- Hemingway: Modeling Distributed Optimization Algorithms
- Asynchronous Stochastic Gradient Descent with Variance Reduction for Non-Convex Optimization
- Local AdaAlter: Communication-Efficient Stochastic Gradient Descent with Adaptive Learning Rates
- Asynchronous Accelerated Proximal Stochastic Gradient for Strongly Convex Distributed Finite Sums
- Building a Fine-Grained Entity Typing System Overnight for a New X (X = Language, Domain, Genre)
- Parallelizing Word2Vec in Shared and Distributed Memory
- Retouchdown: Adding Touchdown to StreetLearn as a Shareable Resource for Language Grounding Tasks in Street View
- Ensuring Rapid Mixing and Low Bias for Asynchronous Gibbs Sampling
- Least Squares Revisited: Scalable Approaches for Multi-class Prediction
- Gradient Diversity: a Key Ingredient for Scalable Distributed Learning
- Dorylus: Affordable, Scalable, and Accurate GNN Training with Distributed CPU Servers and Serverless Threads
- Accuracy-Efficiency Trade-Offs and Accountability in Distributed ML Systems
- Deep Learning at 15PF: Supervised and Semi-Supervised Classification for Scientific Data
- The Potential of the Intel Xeon Phi for Supervised Deep Learning
- Parle: parallelizing stochastic gradient descent
- GT-SEER: Geo-Temporal SEquential Embedding Rank for Point-of-interest Recommendation
- Communication trade-offs for synchronized distributed SGD with large step size
- MXNET-MPI: Embedding MPI parallelism in Parameter Server Task Model for scaling Deep Learning
- Efficient Parallel Translating Embedding For Knowledge Graphs
- Adaptive learning rates and parallelization for stochastic, sparse, non-smooth gradients
- LOCO: Distributing Ridge Regression with Random Projections
- Towards Crowdsourced Training of Large Neural Networks using Decentralized Mixture-of-Experts
- An Accelerated Decentralized Stochastic Proximal Algorithm for Finite Sums
- Large Scale Kernel Learning using Block Coordinate Descent
- Achieving Linear Convergence in Distributed Asynchronous Multi-agent Optimization
- Backprop with Approximate Activations for Memory-efficient Network Training
- Infrastructure for Usable Machine Learning: The Stanford DAWN Project
- Mapping Instructions to Actions in 3D Environments with Visual Goal Prediction
- Splash: User-friendly Programming Interface for Parallelizing Stochastic Algorithms
- AspEm: Embedding Learning by Aspects in Heterogeneous Information Networks
- Distributed Proximal Gradient Algorithm for Partially Asynchronous Computer Clusters
- Variational consensus Monte Carlo
- DimmWitted: A Study of Main-Memory Statistical Analytics
- Declarative Machine Learning - A Classification of Basic Properties and Types
- Optimistic Concurrency Control for Distributed Unsupervised Learning
- Matrix Completion under Interval Uncertainty
- Exploring the Design Space of Deep Convolutional Neural Networks at Large Scale
- Distributed Learning of Deep Neural Networks using Independent Subnet Training
- Taming Momentum in a Distributed Asynchronous Environment
- OmniNet: A unified architecture for multi-modal multi-task learning
- Tupleware: Redefining Modern Analytics
- DeepSpark: A Spark-Based Distributed Deep Learning Framework for Commodity Clusters
- Orchestrating the Development Lifecycle of Machine Learning-Based IoT Applications: A Taxonomy and Survey
- The Convergence of Sparsified Gradient Methods
- CROSSBOW: Scaling Deep Learning with Small Batch Sizes on Multi-GPU Servers
- FedDR -- Randomized Douglas-Rachford Splitting Algorithms for Nonconvex Federated Composite Optimization
- Big Learning with Bayesian Methods
- Strategies and Principles of Distributed Machine Learning on Big Data
- rTop-k: A Statistical Estimation Approach to Distributed SGD
- Domain-specific Communication Optimization for Distributed DNN Training
- Understanding and Optimizing the Performance of Distributed Machine Learning Applications on Apache Spark
- Communication-Efficient Distributed Optimization in Networks with Gradient Tracking and Variance Reduction
- Parity Models: A General Framework for Coding-Based Resilience in ML Inference
- SparCML: High-Performance Sparse Communication for Machine Learning
- Vector representations of text data in deep learning
- CuMF_SGD: Fast and Scalable Matrix Factorization
- Ripple: A Practical Declarative Programming Framework for Serverless Compute
- Characterizing Impacts of Heterogeneity in Federated Learning upon Large-Scale Smartphone Data
- Large-scale randomized-coordinate descent methods with non-separable linear constraints
- Stochastic Block Mirror Descent Methods for Nonsmooth and Stochastic Optimization
- GradiVeQ: Vector Quantization for Bandwidth-Efficient Gradient Aggregation in Distributed CNN Training
- Learning Multivariate Hawkes Processes at Scale
- Asynchronous Stochastic Proximal Optimization Algorithms with Variance Reduction
- Byzantine Resilient Non-Convex SVRG with Distributed Batch Gradient Computations
- Asynchronous Complex Analytics in a Distributed Dataflow Architecture
- The End of Slow Networks: It's Time for a Redesign
- KAISA: An Adaptive Second-Order Optimizer Framework for Deep Neural Networks
- Zeroth-order Asynchronous Doubly Stochastic Algorithm with Variance Reduction
- Block-diagonal Hessian-free Optimization for Training Neural Networks
- RankMap: A Platform-Aware Framework for Distributed Learning from Dense Datasets
- Fast Differentially Private Matrix Factorization
- Parallel Training of Deep Networks with Local Updates
- Learning Emoji Embeddings using Emoji Co-occurrence Network Graph
- Nonasymptotic convergence of stochastic proximal point algorithms for constrained convex optimization
- Amortized Analysis on Asynchronous Gradient Descent
- Petuum: A New Platform for Distributed Machine Learning on Big Data
- PRETZEL: Opening the Black Box of Machine Learning Prediction Serving Systems
- Comparing Deep Reinforcement Learning and Evolutionary Methods in Continuous Control
- Keeping CALM: When Distributed Consistency is Easy
- Distributed Stochastic Algorithms for High-rate Streaming Principal Component Analysis
- Robust and Communication-Efficient Collaborative Learning
- Large-Scale Distributed Bayesian Matrix Factorization using Stochastic Gradient MCMC
- Stochastic gradient descent methods for estimation with large data sets
- Communication Optimality Trade-offs For Distributed Estimation
- SMART: The Stochastic Monotone Aggregated Root-Finding Algorithm
- Distributed Training Large-Scale Deep Architectures
- Asynchronous Stochastic Coordinate Descent: Parallelism and Convergence Properties
- Deep Probabilistic Programming Languages: A Qualitative Study
- Approximate Decentralized Bayesian Inference
- Distributed Deep Learning with Event-Triggered Communication
- Parallelizing Word2Vec in Multi-Core and Many-Core Architectures
- Asynchronous Parallel Stochastic Gradient Descent - A Numeric Core for Scalable Distributed Machine Learning Algorithms
- Heterogeneity-Aware Asynchronous Decentralized Training
- Faster Neural Network Training with Approximate Tensor Operations
- Asynchronous Stochastic Block Coordinate Descent with Variance Reduction
- EventGraD: Event-Triggered Communication in Parallel Machine Learning
- SMORE: Knowledge Graph Completion and Multi-hop Reasoning in Massive Knowledge Graphs
- Stochastic Gradient Descent on Highly-Parallel Architectures
- Structure Regularization for Structured Prediction: Theories and Experiments
- Training Federated GANs with Theoretical Guarantees: A Universal Aggregation Approach
- Distributed Sketching Methods for Privacy Preserving Regression
- Elastic Consistency: A General Consistency Model for Distributed Stochastic Gradient Descent
- Beyond Human-Level Accuracy: Computational Challenges in Deep Learning
- At Stability's Edge: How to Adjust Hyperparameters to Preserve Minima Selection in Asynchronous Training of Neural Networks?
- Distributed Inexact Damped Newton Method: Data Partitioning and Load-Balancing
- Weld: Rethinking the Interface Between Data-Intensive Applications
- Variance Reduction for Distributed Stochastic Gradient Descent
- Anarchic Federated Learning
- A Block-wise, Asynchronous and Distributed ADMM Algorithm for General Form Consensus Optimization
- A Hybrid Method of Combinatorial Search and Coordinate Descent for Discrete Optimization
- ParaGraphE: A Library for Parallel Knowledge Graph Embedding
- Distributed Stochastic Optimization of the Regularized Risk
- Fast, Accurate, and Scalable Method for Sparse Coupled Matrix-Tensor Factorization
- SGD for Structured Nonconvex Functions: Learning Rates, Minibatching and Interpolation
- Random Shuffling Beats SGD after Finite Epochs
- Model-Parallel Inference for Big Topic Models
- Nondeterminism and Instability in Neural Network Optimization
- Collectively Embedding Multi-Relational Data for Predicting User Preferences
- A Unified Analysis of Stochastic Optimization Methods Using Jump System Theory and Quadratic Constraints
- DBS: Dynamic Batch Size For Distributed Deep Neural Network Training
- Helix: Holistic Optimization for Accelerating Iterative Machine Learning
- A machine-compiled macroevolutionary history of Phanerozoic life
- Partitioning Data on Features or Samples in Communication-Efficient Distributed Optimization?
- Optimizing Multi-GPU Parallelization Strategies for Deep Learning Training
- Faster and Cheaper: Parallelizing Large-Scale Matrix Factorization on GPUs
- When FastText Pays Attention: Efficient Estimation of Word Representations using Constrained Positional Weighting
- Nested Dithered Quantization for Communication Reduction in Distributed Training
- Fast Asynchronous Parallel Stochastic Gradient Decent
- Do optimization methods in deep learning applications matter?
- Efficient Communications in Training Large Scale Neural Networks
- Taming Convergence for Asynchronous Stochastic Gradient Descent with Unbounded Delay in Non-Convex Learning
- Async-RED: A Provably Convergent Asynchronous Block Parallel Stochastic Method using Deep Denoising Priors
- Distributed Asynchronous Dual Free Stochastic Dual Coordinate Ascent
- TMAC: A Toolbox of Modern Async-Parallel, Coordinate, Splitting, and Stochastic Methods
- DUAL-LOCO: Distributing Statistical Estimation Using Random Projections
- Accelerated, Optimal, and Parallel: Some Results on Model-Based Stochastic Optimization
- The Three Pillars of Machine Programming
- Scalable Parallel Factorizations of SDD Matrices and Efficient Sampling for Gaussian Graphical Models
- Probabilistic Synchronous Parallel
- Towards a Unified Architecture for in-RDBMS Analytics
- DS-MLR: Exploiting Double Separability for Scaling up Distributed Multinomial Logistic Regression
- Training Neural Networks with Fixed Sparse Masks
- Para-active learning
- Trends and Advancements in Deep Neural Network Communication
- Distributed Deep Learning in Open Collaborations
- Adaptive Sampling Distributed Stochastic Variance Reduced Gradient for Heterogeneous Distributed Datasets
- ASAP: Asynchronous Approximate Data-Parallel Computation
- Dynamic Stale Synchronous Parallel Distributed Training for Deep Learning
- KeystoneML: Optimizing Pipelines for Large-Scale Advanced Analytics
- Decoupled Asynchronous Proximal Stochastic Gradient Descent with Variance Reduction
- Elastic Gossip: Distributing Neural Network Training Using Gossip-like Protocols
- Asynchronous Parallel Empirical Variance Guided Algorithms for the Thresholding Bandit Problem
- Vertex-Context Sampling for Weighted Network Embedding
- Deep Determinantal Point Processes
- Accelerating Asynchronous Stochastic Gradient Descent for Neural Machine Translation
- Optimal Mini-Batch Size Selection for Fast Gradient Descent
- Make Workers Work Harder: Decoupled Asynchronous Proximal Stochastic Gradient Descent
- ParMAC: distributed optimisation of nested functions, with application to learning binary autoencoders
- On Unbounded Delays in Asynchronous Parallel Fixed-Point Algorithms
- Asynchronous Stochastic Proximal Methods for Nonconvex Nonsmooth Optimization
- Graph Embeddings at Scale
- Layered gradient accumulation and modular pipeline parallelism: fast and efficient training of large language models
- Parallel and Flexible Sampling from Autoregressive Models via Langevin Dynamics
- Where Is the Normative Proof? Assumptions and Contradictions in ML Fairness Research
- A Provably Communication-Efficient Asynchronous Distributed Inference Method for Convex and Nonconvex Problems
- Federated Doubly Stochastic Kernel Learning for Vertically Partitioned Data
- Communication-Efficient Policy Gradient Methods for Distributed Reinforcement Learning
- Anytime Stochastic Gradient Descent: A Time to Hear from all the Workers
- Triple2Vec: Learning Triple Embeddings from Knowledge Graphs
- Network-accelerated Distributed Machine Learning Using MLFabric
- Transfer Deep Learning for Low-Resource Chinese Word Segmentation with a Novel Neural Network
- Deep Learning Approaches for Image Retrieval and Pattern Spotting in Ancient Documents
- Near-Data Processing for Differentiable Machine Learning Models
- Parallel Stochastic Gradient Descent with Sound Combiners
- Extreme Stochastic Variational Inference: Distributed and Asynchronous
- Efficient Parallel Learning of Word2Vec
- Coordinate Descent Algorithms
- Blazes: Coordination Analysis for Distributed Programs
- Parallel and Distributed Collaborative Filtering: A Survey
- Bregman Monotone Operator Splitting
- Streaming Principal Component Analysis From Incomplete Data
- Net2Vec: Deep Learning for the Network
- The Convergence of Stochastic Gradient Descent in Asynchronous Shared Memory
- A Class of Parallel Doubly Stochastic Algorithms for Large-Scale Learning
- Differential Equations for Modeling Asynchronous Algorithms
- Asynchronous Decentralized Stochastic Optimization in Heterogeneous Networks
- Parallel Coordinate Descent Newton Method for Efficient -Regularized Minimization
- The Scalability for Parallel Machine Learning Training Algorithm: Dataset Matters
- Anomaly Detection and Correction in Large Labeled Bipartite Graphs
- A Survey on Large-scale Machine Learning
- Cataloging the Visible Universe through Bayesian Inference at Petascale
- LAGC: Lazily Aggregated Gradient Coding for Straggler-Tolerant and Communication-Efficient Distributed Learning
- Asynchronous Distributed Optimization with Redundancy in Cost Functions
- Accelerated Sparsified SGD with Error Feedback
- Parallel and Communication Avoiding Least Angle Regression
- GEVO: GPU Code Optimization using Evolutionary Computation
- Heterogeneous CPU+GPU Stochastic Gradient Descent Algorithms
- A Unifying Framework for Variance Reduction Algorithms for Finding Zeroes of Monotone Operators
- Making Asynchronous Stochastic Gradient Descent Work for Transformers
- Sync-Switch: Hybrid Parameter Synchronization for Distributed Deep Learning
- Consistent Lock-free Parallel Stochastic Gradient Descent for Fast and Stable Convergence
- Relaxed Scheduling for Scalable Belief Propagation
- Parallel and distributed asynchronous adaptive stochastic gradient methods
- Asynchrony and Acceleration in Gossip Algorithms
- Adaptive Elastic Training for Sparse Deep Learning on Heterogeneous Multi-GPU Servers
- Hogwild! over Distributed Local Data Sets with Linearly Increasing Mini-Batch Sizes
- Reified Context Models
- Communication-Efficient Distributed Optimization with Quantized Preconditioners
- Matrix Completion via Factorizing Polynomials
- Reframing Threat Detection: Inside esINSIDER
- F10-SGD: Fast Training of Elastic-net Linear Models for Text Classification and Named-entity Recognition
- Distributed Low-rank Subspace Segmentation
- Sparsification as a Remedy for Staleness in Distributed Asynchronous SGD
- Adaptive In-network Collaborative Caching for Enhanced Ensemble Deep Learning at Edge
- Asynchronous Stochastic Optimization Robust to Arbitrary Delays
- Avoiding communication in primal and dual block coordinate descent methods
- Graph Balancing for Distributed Subgradient Methods over Directed Graphs
- Stochastic Gradient MCMC with Stale Gradients
- Node Embedding via Word Embedding for Network Community Discovery
- An Approximate, Efficient Solver for LP Rounding
- A Block Decomposition Algorithm for Sparse Optimization
- Block Distributed Majorize-Minimize Memory Gradient Algorithm and its application to 3D image restoration
- Error Compensated Loopless SVRG, Quartz, and SDCA for Distributed Optimization
- Asynchronous Distributed Optimization with Stochastic Delays
- CATERPILLAR: Coarse Grain Reconfigurable Architecture for Accelerating the Training of Deep Neural Networks
- Integrating Deep Learning in Domain Sciences at Exascale
- Simple and Efficient Parallelization for Probabilistic Temporal Tensor Factorization
- High Throughput Synchronous Distributed Stochastic Gradient Descent
- Compressed Coded Distributed Computing
- Improving Skip-Gram based Graph Embeddings via Centrality-Weighted Sampling
- Distributed deep learning on edge-devices: feasibility via adaptive compression
- ASYMP: Fault-tolerant Mining of Massive Graphs
- A Generic Online Parallel Learning Framework for Large Margin Models
- Pushing the boundaries of parallel Deep Learning -- A practical approach
- Distributed stochastic optimization for deep learning (thesis)
- Adversarial Delays in Online Strongly-Convex Optimization
- A Sparse Completely Positive Relaxation of the Modularity Maximization for Community Detection
- Finite-Time Consensus Learning for Decentralized Optimization with Nonlinear Gossiping
- ANDRUSPEX : Leveraging Graph Representation Learning to Predict Harmful App Installations on Mobile Devices
- Gradient Sparification for Asynchronous Distributed Training
- Weighted parallel SGD for distributed unbalanced-workload training system
- Model Size Reduction Using Frequency Based Double Hashing for Recommender Systems
- Learning a Predictive Model for Music Using PULSE
- Communication-Efficient Asynchronous Stochastic Frank-Wolfe over Nuclear-norm Balls
- Deep Learning At Scale and At Ease
- Secure Distributed Training at Scale
- A Systematic Investigation of KB-Text Embedding Alignment at Scale
- A Sharp Convergence Rate for the Asynchronous Stochastic Gradient Descent
- A Parallel and Efficient Algorithm for Learning to Match
- A Forest Mixture Bound for Block-Free Parallel Inference
- Tell Me Something New: A New Framework for Asynchronous Parallel Learning
- FULL-W2V: Fully Exploiting Data Reuse for W2V on GPU-Accelerated Systems
- Matrix Factorization on GPUs with Memory Optimization and Approximate Computing
- GaDei: On Scale-up Training As A Service For Deep Learning
- A Model Parallel Proximal Stochastic Gradient Algorithm for Partially Asynchronous Systems
- Fully Asynchronous Stochastic Coordinate Descent: A Tight Lower Bound on the Parallelism Achieving Linear Speedup
- Distributed Networked Real-time Learning
- Homomorphic Parameter Compression for Distributed Deep Learning Training
- Approaching the Ad Placement Problem with Online Linear Classification: The winning solution to the NIPS'17 Ad Placement Challenge
- Optimization for Supervised Machine Learning: Randomized Algorithms for Data and Parameters
- Fully Distributed and Asynchronized Stochastic Gradient Descent for Networked Systems
- Accelerating Perturbed Stochastic Iterates in Asynchronous Lock-Free Optimization
- Efficient Matrix Factorization on Heterogeneous CPU-GPU Systems
- IS-ASGD: Accelerating Asynchronous SGD using Importance Sampling
- Consensus-Based Modelling using Distributed Feature Construction
- Skewness Ranking Optimization for Personalized Recommendation
- Online Learning with Optimism and Delay
- Balancing the Communication Load of Asynchronously Parallelized Machine Learning Algorithms
- Associative Convolutional Layers
- Reducing Data Motion to Accelerate the Training of Deep Neural Networks
- Topic Modeling via Full Dependence Mixtures
- Asynchronous Stochastic Gradient MCMC with Elastic Coupling
- Probabilistic Bag-Of-Hyperlinks Model for Entity Linking
- Partitioning Large Scale Deep Belief Networks Using Dropout
- Practical Precoding via Asynchronous Stochastic Successive Convex Approximation
- Parameter Database : Data-centric Synchronization for Scalable Machine Learning
- Adversarial Training Methods for Network Embedding
- A Random Gossip BMUF Process for Neural Language Modeling
- Large-Scale Stochastic Learning using GPUs
- SCAR: Strong Consistency using Asynchronous Replication with Minimal Coordination
- Oscars: Adaptive Semi-Synchronous Parallel Model for Distributed Deep Learning with Global View
- Collaborative Similarity Embedding for Recommender Systems
- Sketching Linear Classifiers over Data Streams
- Flexible numerical optimization with ensmallen
- Parallel training of linear models without compromising convergence
- Distributed Learning and its Application for Time-Series Prediction