2 papers
cs.DC2025
Deadline-Aware Online Scheduling for LLM Fine-Tuning with Spot Market Predictions
Linggao Kong, Yuedong Xu, Lei Jiao +1
As foundation models grow in size, fine-tuning them becomes increasingly expensive. While GPU spot instances offer a low-cost alternative to on-demand resources, their volatile pri…
cs.DC2025
Diving into 3D Parallelism with Heterogeneous Spot Instance GPUs: Design and Implications
Yuxiao Wang, Yuedong Xu, Qingyang Duan +4
The rapid growth of large language models (LLMs) and the continuous release of new GPU products have significantly increased the demand for distributed training across heterogeneou…