3 papers
cs.DC2025
FUSCO: High-Performance Distributed Data Shuffling via Transformation-Communication Fusion
Zhuoran Zhu, Chunyang Zhu, Hao Lin +9
Large-scale Mixture-of-Experts (MoE) models rely on \emph{expert parallelism} for efficient training and inference, which splits experts across devices and necessitates distributed…
cs.AR2025
A Systematic Characterization of LLM Inference on GPUs
Haonan Wang, Xuxin Xiao, Mingyu Yan +8
This work presents a systematic characterization of Large Language Model (LLM) inference to address fragmented understanding. Through comprehensive experiments, we establish a four…
cs.DC2025
Understanding the Landscape of Ampere GPU Memory Errors
Zhu Zhu, Yu Sun, Dhatri Parakal +9
Graphics Processing Units (GPUs) have become a de facto solution for accelerating high-performance computing (HPC) applications. Understanding their memory error behavior is an ess…