2 papers
cs.PF2018
Hoard: A Distributed Data Caching System to Accelerate Deep Learning Training on the Cloud
Christian Pinto, Yiannis Gkoufas, Andrea Reale +2
Deep Learning system architects strive to design a balanced system where the computational accelerator -- FPGA, GPU, etc, is not starved for data. Feeding training data fast enough…
cs.NE2016
dMath: A Scalable Linear Algebra and Math Library for Heterogeneous GP-GPU Architectures
Steven Eliuk, Cameron Upright, Anthony Skjellum
A new scalable parallel math library, dMath, is presented in this paper that demonstrates leading scaling when using intranode, or internode, hybrid-parallelism for deep-learning.…