4 papers
HIERA: Workload-Aware Planning Across Implementation Spaces for GPU Kernel Optimization
Jinghao Wang, Qiqi Gu, Chenpeng Wu +3
High-performance GPU kernels underpin modern deep learning and scientific computing. As workloads become increasingly diverse and GPU hardware evolves rapidly, developing efficient…
Do We Need Tensor Cores for Stencil Computations?
Qiqi Gu, Chenpeng Wu, Heng Shi +2
Stencil computation constitutes a cornerstone of scientific computing, serving as a critical kernel in domains ranging from fluid dynamics to weather simulation. While stencil comp…
SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided Swapping
Qiqi GU, Chenpeng Wu, Heng Shi +1
Recent research has focused on accelerating stencil computations by exploiting emerging hardware like Tensor Cores. To leverage these accelerators, the stencil operation must be tr…
Samoyeds: Accelerating MoE Models with Structured Sparsity Leveraging Sparse Tensor Cores
Chenpeng Wu, Qiqi Gu, Heng Shi +2
The escalating size of Mixture-of-Experts (MoE) based Large Language Models (LLMs) presents significant computational and memory challenges, necessitating innovative solutions to e…