2 papers
cs.DC2025
Libra: Unleashing GPU Heterogeneity for High-Performance Sparse Matrix Multiplication
Jinliang Shi, Shigang Li, Youxuan Xu +4
Sparse matrix multiplication operators (i.e., SpMM and SDDMM) are widely used in deep learning and scientific computing. Modern accelerators are commonly equipped with Tensor Core…
cs.DC2025
SparkAttention: High-Performance Multi-Head Attention for Large Models on Volta GPU Architecture
Youxuan Xu, Tong Wu, Shigang Li +2
Transformer are widely used in various fields such as natural language processing and computer vision. However, the training time for large Transformer models can be challenging du…