3 papers
cs.DC2026
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
Chen Zhuang, Lingqi Zhang, Benjamin Brock +5
Distributed Sparse Matrix-Matrix Multiplication (SpMM) is a fundamental operation in high-performance computing and deep learning applications. The major performance bottleneck in…
cs.DC2025
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
Chen Zhuang, Lingqi Zhang, Du Wu +8
Graph Convolutional Networks (GCNs), particularly for large-scale graphs, are crucial across numerous domains. However, training distributed full-batch GCNs on large-scale graphs s…
cs.DC2025
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
Lingqi Zhang, Jiajun Huang, Sheng Di +2
Tensor cores are specialized processing units within GPUs that have demonstrated significant efficiency gains in compute-bound applications such as Deep Learning Training by accele…