176 citations · 296 across the 16 of their papers we have counts for
10 papers · 1 filter
HEAT: A Highly Efficient and Affordable Training System for Collaborative Filtering Based Recommendation on CPUs
Chengming Zhang, Shaden Smith, Baixi Sun +6
Collaborative filtering (CF) has been proven to be one of the most effective techniques for recommendation. Among all CF approaches, SimpleX is the state-of-the-art method that ado…
MSREP: A Fast yet Light Sparse Matrix Framework for Multi-GPU Systems
Jieyang Chen, Chenhao Xie, Jesun S Firoz +5
Sparse linear algebra kernels play a critical role in numerous applications, covering from exascale scientific simulation to large-scale data analytics. Offloading linear algebra k…
Towards Efficient Architecture and Algorithms for Sensor Fusion
Zhendong Wang, Xiaoming Zeng, Shuaiwen Leon Song +1
The safety of an automated vehicle hinges crucially upon the accuracy of perception and decision-making latency. Under these stringent requirements, future automated cars are usual…
MAPA: Multi-Accelerator Pattern Allocation Policy for Multi-Tenant GPU Servers
Kiran Ranganath, Joshua D. Suetterlein, Joseph B. Manzano +2
Multi-accelerator servers are increasingly being deployed in shared multi-tenant environments (such as in cloud data centers) in order to meet the demands of large-scale compute-in…
Fast and Scalable Sparse Triangular Solver for Multi-GPU Based HPC Architectures
Chenhao Xie, Jieyang Chen, Jesun S Firoz +5
Designing efficient and scalable sparse linear algebra kernels on modern multi-GPU based HPC systems is a daunting task due to significant irregular memory references and workload…
A Novel Memory-Efficient Deep Learning Training Framework via Error-Bounded Lossy Compression
Sian Jin, Guanpeng Li, Shuaiwen Leon Song +1
Deep neural networks (DNNs) are becoming increasingly deeper, wider, and non-linear due to the growing demands on prediction accuracy and analysis quality. When training a DNN mode…