2 papers
cs.PL2025
A Performance Model for Warp Specialization Kernels
Zhengyang Liu, Vinod Grover
This paper presents a performance model tailored for warp specialization kernels, focusing on factors such as warp size, tilling size, input matrix size, memory bandwidth, and thre…
cs.DC2020
Synthesizing Optimal Collective Algorithms
Zixian Cai, Zhengyang Liu, Saeed Maleki +4
Collective communication algorithms are an important component of distributed computation. Indeed, in the case of deep-learning, collective communication is the Amdahl's bottleneck…