171 citations · 311 across the 13 of their papers we have counts for
28 papers
Distributed Deep Learning in Open Collaborations
Michael Diskin, Alexey Bukhtiyarov, Max Ryabinin +13
Modern deep learning applications require increasingly more compute to train state-of-the-art models. To address this demand, large corporations and institutions use dedicated High…
Horizontally Fused Training Array: An Effective Hardware Utilization Squeezer for Training Novel Deep Learning Models
Shang Wang, Peiming Yang, Yuxuan Zheng +2
Driven by the tremendous effort in researching novel deep learning (DL) algorithms, the training cost of developing new models increases staggeringly in recent years. We analyze GP…
RL-Scope: Cross-Stack Profiling for Deep Reinforcement Learning Workloads
James Gleeson, Srivatsan Krishnan, Moshe Gabel +3
Deep reinforcement learning (RL) has made groundbreaking advancements in robotics, data center management and other applications. Unfortunately, system-level bottlenecks in RL work…
A Runtime-Based Computational Performance Predictor for Deep Neural Network Training
Geoffrey X. Yu, Yubo Gao, Pavel Golikov +1
Deep learning researchers and practitioners usually leverage GPUs to help train their deep neural networks (DNNs) faster. However, choosing which GPU to use is challenging both bec…
LifeStream: A High-Performance Stream Processing Engine for Periodic Streams
Anand Jayarajan, Kimberly Hau, Andrew Goodwin +1
Hospitals around the world collect massive amounts of physiological data from their patients every day. Recently, there has been an increase in research interest to subject this da…
IOS: Inter-Operator Scheduler for CNN Acceleration
Yaoyao Ding, Ligeng Zhu, Zhihao Jia +2
To accelerate CNN inference, existing deep learning frameworks focus on optimizing intra-operator parallelization. However, a single operator can no longer fully utilize the availa…