3 papers
cs.DC2026
HIERA: Workload-Aware Planning Across Implementation Spaces for GPU Kernel Optimization
Jinghao Wang, Qiqi Gu, Chenpeng Wu +3
High-performance GPU kernels underpin modern deep learning and scientific computing. As workloads become increasingly diverse and GPU hardware evolves rapidly, developing efficient…
cs.DC2026
Do We Need Tensor Cores for Stencil Computations?
Qiqi Gu, Chenpeng Wu, Heng Shi +2
Stencil computation constitutes a cornerstone of scientific computing, serving as a critical kernel in domains ranging from fluid dynamics to weather simulation. While stencil comp…
cs.DC2025
SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided Swapping
Qiqi GU, Chenpeng Wu, Heng Shi +1
Recent research has focused on accelerating stencil computations by exploiting emerging hardware like Tensor Cores. To leverage these accelerators, the stencil operation must be tr…