2 papers
cs.DC2025
Zorse: Optimizing LLM Training Efficiency on Heterogeneous GPU Clusters
Runsheng Benson Guo, Utkarsh Anand, Khuzaima Daudjee +1
Large language models (LLMs) require vast amounts of GPU compute to train, but limited availability and high costs of GPUs make homogeneous clusters impractical for many organizati…
cs.DC2025
FreeRide: Harvesting Bubbles in Pipeline Parallelism
Jiashu Zhang, Zihan Pan, Molly +3
The occurrence of bubbles in pipeline parallelism is an inherent limitation that can account for more than 40% of the large language model (LLM) training time and is one of the mai…