9 citations · 10 across the 5 of their papers we have counts for
Showing cs.DCShow all
3 papers · 1 filter
cs.DC2025
ElasWave: An Elastic-Native System for Scalable Hybrid-Parallel Training
Xueze Kang, Guangyu Xiang, Yuxin Wang +16
Large-scale LLM pretraining now runs across -- accelerators, making failures routine and elasticity mandatory. We posit that an elastic-native training system must join…
cs.DC2024★ 9 cited
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
Yuxin Wang, Yuhan Chen, Zeyu Li +11
Serving systems for Large Language Models (LLMs) are often optimized to improve quality of service (QoS) and throughput. However, due to the lack of open-source LLM serving workloa…
cs.DC2023★ 1 cited
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
Yuxin Wang, Xueze Kang, Shaohuai Shi +8
To efficiently scale large model (LM) training, researchers transition from data parallelism (DP) to hybrid parallelism (HP) on GPU clusters, which frequently experience hardware a…