176 citations · 279 across the 12 of their papers we have counts for
16 papers
MSREP: A Fast yet Light Sparse Matrix Framework for Multi-GPU Systems
Jieyang Chen, Chenhao Xie, Jesun S Firoz +5
Sparse linear algebra kernels play a critical role in numerous applications, covering from exascale scientific simulation to large-scale data analytics. Offloading linear algebra k…
Towards Efficient Architecture and Algorithms for Sensor Fusion
Zhendong Wang, Xiaoming Zeng, Shuaiwen Leon Song +1
The safety of an automated vehicle hinges crucially upon the accuracy of perception and decision-making latency. Under these stringent requirements, future automated cars are usual…
Shift-BNN: Highly-Efficient Probabilistic Bayesian Neural Network Training via Memory-Friendly Pattern Retrieving
Qiyu Wan, Haojun Xia, Xingyao Zhang +3
Bayesian Neural Networks (BNNs) that possess a property of uncertainty estimation have been increasingly adopted in a wide range of safety-critical AI applications which demand rel…
MAPA: Multi-Accelerator Pattern Allocation Policy for Multi-Tenant GPU Servers
Kiran Ranganath, Joshua D. Suetterlein, Joseph B. Manzano +2
Multi-accelerator servers are increasingly being deployed in shared multi-tenant environments (such as in cloud data centers) in order to meet the demands of large-scale compute-in…
Dr. Top-k: Delegate-Centric Top-k on GPUs
Anil Gaihre, Da Zheng, Scott Weitze +5
Recent top- computation efforts explore the possibility of revising various sorting algorithms to answer top- queries on GPUs. These endeavors, unfortunately, perform signifi…
Randomness In Neural Network Training: Characterizing The Impact of Tooling
Donglin Zhuang, Xingyao Zhang, Shuaiwen Leon Song +1
The quest for determinism in machine learning has disproportionately focused on characterizing the impact of noise introduced by algorithmic design choices. In this work, we addres…