176 citations · 207 across the 11 of their papers we have counts for
9 papers · 1 filter
MSREP: A Fast yet Light Sparse Matrix Framework for Multi-GPU Systems
Jieyang Chen, Chenhao Xie, Jesun S Firoz +5
Sparse linear algebra kernels play a critical role in numerous applications, covering from exascale scientific simulation to large-scale data analytics. Offloading linear algebra k…
Fast and Scalable Sparse Triangular Solver for Multi-GPU Based HPC Architectures
Chenhao Xie, Jieyang Chen, Jesun S Firoz +5
Designing efficient and scalable sparse linear algebra kernels on modern multi-GPU based HPC systems is a daunting task due to significant irregular memory references and workload…
ARENA: Asynchronous Reconfigurable Accelerator Ring to Enable Data-Centric Parallel Computing
Cheng Tan, Chenhao Xie, Tong Geng +4
The next generation HPC and data centers are likely to be reconfigurable and data-centric due to the trend of hardware specialization and the emergence of data-driven applications.…
Accelerating Binarized Neural Networks via Bit-Tensor-Cores in Turing GPUs
Ang Li, Simon Su
Despite foreseeing tremendous speedups over conventional deep neural networks, the performance advantage of binarized neural networks (BNNs) has merely been showcased on general-pu…
CSB-RNN: A Faster-than-Realtime RNN Acceleration Framework with Compressed Structured Blocks
Runbin Shi, Peiyan Dong, Tong Geng +6
Recurrent neural networks (RNNs) have been widely adopted in temporal sequence analysis, where realtime performance is often in demand. However, RNNs suffer from heavy computationa…
A Parallel Sparse Tensor Benchmark Suite on CPUs and GPUs
Jiajia Li, Mahesh Lakshminarasimhan, Xiaolong Wu +3
Tensor computations present significant performance challenges that impact a wide spectrum of applications ranging from machine learning, healthcare analytics, social network analy…