2 papers
cs.AR2026
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
Euijun Chung, Yuxiao Jia, Aaron Jezghani +1
Large-scale machine learning workloads increasingly rely on multi-GPU systems, yet their performance is often limited by an overlooked component: the CPU. Through a detailed study…
cs.DC2025
Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal Perspective
Seokjin Go, Joongun Park, Spandan More +5
The rapid scaling of Large Language Models (LLMs) has pushed training workloads far beyond the limits of single-node analysis, demanding a deeper understanding of how these models…