177 citations · 182 across the 3 of their papers we have counts for
5 papers
Heterogeneity-Aware Asynchronous Decentralized Training
Qinyi Luo, Jiaao He, Youwei Zhuo +1
Distributed deep learning training usually adopts All-Reduce as the synchronization mechanism for data parallel algorithms due to its high performance in homogeneous environment. H…
HyPar: Towards Hybrid Parallelism for Deep Learning Accelerator Array
Linghao Song, Jiachen Mao, Youwei Zhuo +3
With the rise of artificial intelligence in recent years, Deep Neural Networks (DNNs) have been widely used in many domains. To achieve high performance and energy efficiency, hard…
E-RNN: Design Optimization for Efficient Recurrent Neural Networks in FPGAs
Zhe Li, Caiwen Ding, Siyue Wang +8
Recurrent Neural Networks (RNNs) are becoming increasingly important for time series-related applications which require efficient and real-time implementations. The two major types…
CirCNN: Accelerating and Compressing Deep Neural Networks Using Block-CirculantWeight Matrices
Caiwen Ding, Siyu Liao, Yanzhi Wang +13
Large-scale deep neural networks (DNNs) are both compute and memory intensive. As the size of DNNs continues to grow, it is critical to improve the energy efficiency and performanc…
GraphR: Accelerating Graph Processing Using ReRAM
Linghao Song, Youwei Zhuo, Xuehai Qian +2
This paper presents GRAPHR, the first ReRAM-based graph processing accelerator. GRAPHR follows the principle of near-data processing and explores the opportunity of performing mass…