109 citations · 116 across the 13 of their papers we have counts for
3 papers · 1 filter
mKernel: Fast Multi-GPU, Multi-Node Fused Kernels
Ziming Mao, Yihan Zhang, Shawn Wei Chew +5
Communication has become a bottleneck in distributed training and inference of large models. Overlapping communication with computation at the granularity of kernels, on separate s…
UCCL-EP: Portable Expert-Parallel Communication
Ziming Mao, Yihan Zhang, Chihan Cui +9
Mixture-of-Experts (MoE) workloads rely on expert parallelism (EP) to achieve high GPU efficiency. State-of-the-art EP communication systems such as DeepEP demonstrate strong perfo…
I've Got 99 Problems But FLOPS Ain't One
Alexandru M. Gherghescu, Vlad-Andrei Bădoiu, Alexandru Agache +4
Hyperscalers dominate the landscape of large network deployments, yet they rarely share data or insights about the challenges they face. In light of this supremacy, what problems c…