SparkNet: Training Deep Networks in Spark
arXiv:1511.06051
Abstract
Training deep networks is a time-consuming process, with networks for object recognition often requiring multiple days to train. For this reason, leveraging the resources of a cluster to speed up training is an important area of work. However, widely-popular batch-processing computational frameworks like MapReduce and Spark were not designed to support the asynchronous and communication-intensive workloads of existing distributed deep learning systems. We introduce SparkNet, a framework for training deep networks in Spark. Our implementation includes a convenient interface for reading data from Spark RDDs, a Scala interface to the Caffe deep learning framework, and a lightweight multi-dimensional tensor library. Using a simple parallelization scheme for stochastic gradient descent, SparkNet scales well with the cluster size and tolerates very high-latency communication. Furthermore, it is easy to deploy and use with no parameter tuning, and it is compatible with existing Caffe models. We quantify the dependence of the speedup obtained by SparkNet on the number of machines, the communication frequency, and the cluster's communication overhead, and we benchmark our system's performance on the ImageNet dataset.
12 pages, 7 figures
References in corpus (3)
Cited by in corpus (24)
- TensorFlow: A system for large-scale machine learning
- Convergence of Edge Computing and Deep Learning: A Comprehensive Survey
- Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training
- A Field Guide to Federated Optimization
- Cooperative SGD: A unified Framework for the Design and Analysis of Communication-Efficient SGD Algorithms
- Deep Learning in the Automotive Industry: Applications and Tools
- Distributed Statistical Machine Learning in Adversarial Settings: Byzantine Gradient Descent
- GRNN: Generative Regression Neural Network -- A Data Leakage Attack for Federated Learning
- GossipGraD: Scalable Deep Learning using Gossip Communication based Asynchronous Gradient Descent
- Characterizing Deep-Learning I/O Workloads in TensorFlow
- Omnivore: An Optimizer for Multi-device Deep Learning on CPUs and GPUs
- Poseidon: A System Architecture for Efficient GPU-based Deep Learning on Multiple Machines
- Eavesdrop the Composition Proportion of Training Labels in Federated Learning
- Adaptive Communication Strategies to Achieve the Best Error-Runtime Trade-off in Local-Update SGD
- TF-Replicator: Distributed Machine Learning for Researchers
- Parle: parallelizing stochastic gradient descent
- Exploring the Design Space of Deep Convolutional Neural Networks at Large Scale
- Strategies and Principles of Distributed Machine Learning on Big Data
- A Data and Model-Parallel, Distributed and Scalable Framework for Training of Deep Networks in Apache Spark
- SAPAG: A Self-Adaptive Privacy Attack From Gradients
- Tensor Relational Algebra for Machine Learning System Design
- Distributed stochastic optimization for deep learning (thesis)
- OD-SGD: One-step Delay Stochastic Gradient Descent for Distributed Training
- HAlign-II: efficient ultra-large multiple sequence alignment and phylogenetic tree reconstruction with distributed and parallel computing