3 papers
cs.NI2025
A Queueing Theoretic Perspective on Low-Latency LLM Inference with Variable Token Length
Yuqing Yang, Yuedong Xu, Lei Jiao
Large language models (LLMs) propel the prosperity of interactive AI applications showcased by ChatGPT that demand timely response of inference services. However, LLM inference is…
cs.DC2025
Deadline-Aware Online Scheduling for LLM Fine-Tuning with Spot Market Predictions
Linggao Kong, Yuedong Xu, Lei Jiao +1
As foundation models grow in size, fine-tuning them becomes increasingly expensive. While GPU spot instances offer a low-cost alternative to on-demand resources, their volatile pri…
cs.DC2025
Diving into 3D Parallelism with Heterogeneous Spot Instance GPUs: Design and Implications
Yuxiao Wang, Yuedong Xu, Qingyang Duan +4
The rapid growth of large language models (LLMs) and the continuous release of new GPU products have significantly increased the demand for distributed training across heterogeneou…