2 papers
cs.DC2025
Libra: Unleashing GPU Heterogeneity for High-Performance Sparse Matrix Multiplication
Jinliang Shi, Shigang Li, Youxuan Xu +4
Sparse matrix multiplication operators (i.e., SpMM and SDDMM) are widely used in deep learning and scientific computing. Modern accelerators are commonly equipped with Tensor Core…
cs.PF2023
A Performance-Portable SYCL Implementation of CRK-HACC for Exascale
Esteban M. Rangel, S. John Pennycook, Adrian Pope +3
The first generation of exascale systems will include a variety of machine architectures, featuring GPUs from multiple vendors. As a result, many developers are interested in adopt…