3 papers
cs.LG2025
Scorpio: Serving Right Requests at the Right Time for Heterogeneous SLOs in LLM Inference
Yinghao Tang, Tingfeng Lan, Bo Pan +3
Large Language Model (LLM) serving increasingly underpins online Web services such as conversational agents, Web search, and programming assistants, where requests carry heterogene…
cs.LG2023
mLoRA: Fine-Tuning LoRA Adapters via Highly-Efficient Pipeline Parallelism in Multiple GPUs
Zhengmao Ye, Dengchun Li, Zetao Hu +8
Transformer-based, pre-trained large language models (LLMs) have demonstrated outstanding performance across diverse domains, particularly in the emerging {\em pretrain-then-finetu…
cs.DC2023
DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the Cloud
Qinlong Wang, Tingfeng Lan, Yinghao Tang +8
Deep learning recommendation models (DLRM) rely on large embedding tables to manage categorical sparse features. Expanding such embedding tables can significantly enhance model per…