45 citations · 57 across the 3 of their papers we have counts for
4 papers
KAISA: An Adaptive Second-Order Optimizer Framework for Deep Neural Networks
J. Gregory Pauloski, Qi Huang, Lei Huang +4
Kronecker-factored Approximate Curvature (K-FAC) has recently been shown to converge faster in deep neural network (DNN) training than stochastic gradient descent (SGD); however, K…
The Limit of the Batch Size
Yang You, Yuhui Wang, Huan Zhang +3
Large-batch training is an efficient approach for current distributed deep learning systems. It has enabled researchers to reduce the ImageNet/ResNet-50 training from 29 hours to a…
FanStore: Enabling Efficient and Scalable I/O for Distributed Deep Learning
Zhao Zhang, Lei Huang, Uri Manor +5
Emerging Deep Learning (DL) applications introduce heavy I/O workloads on computer clusters. The inherent long lasting, repeated, and random file access pattern can easily saturate…
ImageNet Training in Minutes
Yang You, Zhao Zhang, Cho-Jui Hsieh +2
Finishing 90-epoch ImageNet-1k training with ResNet-50 on a NVIDIA M40 GPU takes 14 days. This training requires 10^18 single precision operations in total. On the other hand, the…