961 citations · 1.1k across the 19 of their papers we have counts for
6 papers · 1 filter
Doing More by Doing Less: How Structured Partial Backpropagation Improves Deep Learning Clusters
Adarsh Kumar, Kausik Subramanian, Shivaram Venkataraman +1
Many organizations employ compute clusters equipped with accelerators such as GPUs and TPUs for training deep learning models in a distributed fashion. Training is resource-intensi…
KAISA: An Adaptive Second-Order Optimizer Framework for Deep Neural Networks
J. Gregory Pauloski, Qi Huang, Lei Huang +4
Kronecker-factored Approximate Curvature (K-FAC) has recently been shown to converge faster in deep neural network (DNN) training than stochastic gradient descent (SGD); however, K…
AutoFreeze: Automatically Freezing Model Blocks to Accelerate Fine-tuning
Yuhan Liu, Saurabh Agarwal, Shivaram Venkataraman
With the rapid adoption of machine learning (ML), a number of domains now use the approach of fine tuning models which were pre-trained on a large corpus of data. However, our expe…
On the Utility of Gradient Compression in Distributed Training Systems
Saurabh Agarwal, Hongyi Wang, Shivaram Venkataraman +1
A rich body of prior work has highlighted the existence of communication bottlenecks in synchronous data-parallel training. To alleviate these bottlenecks, a long line of recent wo…
Accelerating Deep Learning Inference via Learned Caches
Arjun Balasubramanian, Adarsh Kumar, Yuhan Liu +3
Deep Neural Networks (DNNs) are witnessing increased adoption in multiple domains owing to their high accuracy in solving real-world problems. However, this high accuracy has been…
Marius: Learning Massive Graph Embeddings on a Single Machine
Jason Mohoney, Roger Waleffe, Yiheng Xu +2
We propose a new framework for computing the embeddings of large-scale graphs on a single machine. A graph embedding is a fixed length vector representation for each node (and/or e…