507 citations · 661 across the 8 of their papers we have counts for
4 papers · 1 filter
Online Evolutionary Batch Size Orchestration for Scheduling Deep Learning Workloads in GPU Clusters
Zhengda Bian, Shenggui Li, Wei Wang +1
Efficient GPU resource scheduling is essential to maximize resource utilization and save training costs for the increasing amount of deep learning workloads in shared GPU clusters.…
Maximizing Parallelism in Distributed Training for Huge Neural Networks
Zhengda Bian, Qifan Xu, Boxiang Wang +1
The recent Natural Language Processing techniques have been refreshing the state-of-the-art performance at an incredible speed. Training huge language models is therefore an impera…
Accurate, Fast and Scalable Kernel Ridge Regression on Parallel and Distributed Systems
Yang You, James Demmel, Cho-Jui Hsieh +1
We propose two new methods to address the weak scaling problems of KRR: the Balanced KRR (BKRR) and K-means KRR (KKRR). These methods consider alternative ways to partition the inp…
Scaling Deep Learning on GPU and Knights Landing clusters
Yang You, Aydin Buluc, James Demmel
The speed of deep neural networks training has become a big bottleneck of deep learning research and development. For example, training GoogleNet by ImageNet dataset on one Nvidia…