activity
20182022
most citedSuperNeurons: Dynamic GPU Memory Management for Training Deep Neural Networks

176 citations · 207 across the 11 of their papers we have counts for

collaborators

15 papers

cs.DC20222 cited

MSREP: A Fast yet Light Sparse Matrix Framework for Multi-GPU Systems

Jieyang Chen, Chenhao Xie, Jesun S Firoz +5

Sparse linear algebra kernels play a critical role in numerous applications, covering from exascale scientific simulation to large-scale data analytics. Offloading linear algebra k…

cs.AR2022

I-GCN: A Graph Convolutional Network Accelerator with Runtime Locality Enhancement through Islandization

Tong Geng, Chunshu Wu, Yongan Zhang +6

Graph Convolutional Networks (GCNs) have drawn tremendous attention in the past three years. Compared with other deep learning modalities, high-performance hardware acceleration of…

cs.AR20213 cited

G-CoS: GNN-Accelerator Co-Search Towards Both Better Accuracy and Efficiency

Yongan Zhang, Haoran You, Yonggan Fu +3

Graph Neural Networks (GNNs) have emerged as the state-of-the-art (SOTA) method for graph-based learning tasks. However, it still remains prohibitively challenging to inference GNN…

cs.AR20211 cited

Optimizing FPGA-based Accelerator Design for Large-Scale Molecular Similarity Search

Hongwu Peng, Shiyang Chen, Zhepeng Wang +9

Molecular similarity search has been widely used in drug discovery to identify structurally similar compounds from large molecular databases rapidly. With the increasing size of ch…

cs.LG20214 cited

Binary Complex Neural Network Acceleration on FPGA

Hongwu Peng, Shanglin Zhou, Scott Weitze +9

Being able to learn from complex data with phase information is imperative for many signal processing applications. Today' s real-valued deep neural networks (DNNs) have shown effi…

cs.DC2020

Fast and Scalable Sparse Triangular Solver for Multi-GPU Based HPC Architectures

Chenhao Xie, Jieyang Chen, Jesun S Firoz +5

Designing efficient and scalable sparse linear algebra kernels on modern multi-GPU based HPC systems is a daunting task due to significant irregular memory references and workload…