3 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.DC2025
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
Rongzhi Li, Ruogu Du, Zefang Chu +8
Serving Large Language Models (LLMs) is a GPU-intensive task where traditional autoscalers fall short, particularly for modern Prefill-Decode (P/D) disaggregated architectures. Thi…
cs.DC2025
Understanding Stragglers in Large Model Training Using What-if Analysis
Jinkun Lin, Ziheng Jiang, Zuquan Song +13
Large language model (LLM) training is one of the most demanding distributed computations today, often requiring thousands of GPUs with frequent synchronization across machines. Su…
cs.DC2023★ 3 cited
Helios: An Efficient Out-of-core GNN Training System on Terabyte-scale Graphs with In-memory Performance
Jie Sun, Mo Sun, Zheng Zhang +6
Training graph neural networks (GNNs) on large-scale graph data holds immense promise for numerous real-world applications but remains a great challenge. Several disk-based GNN sys…