5 papers
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices
Mingyu Sun, Xiao Zhang, Shen Qu +5
Providing lossless inference services of LLMs on edge devices remains challenging, especially given the extremely tight memory budgets. The existing offloading techniques inevitabl…
Distributed Bilevel Optimization with Dual Pruning for Resource-limited Clients
Mingyi Li, Xiao Zhang, Ruisheng Zheng +4
With the development of large-scale models, traditional distributed bilevel optimization algorithms cannot be applied directly in low-resource clients. The key reason lies in the e…
Data-Free Continual Learning of Server Models in Model-Heterogeneous Cloud-Device Collaboration
Xiao Zhang, Zengzhe Chen, Yuan Yuan +5
The rise of cloud-device collaborative computing has enabled intelligent services to be delivered across distributed edge devices while leveraging centralized cloud resources. In t…
TAMO: Fine-Grained Root Cause Analysis via Tool-Assisted LLM Agent with Multi-Modality Observation Data in Cloud-Native Systems
Xiao Zhang, Qi Wang, Mingyi Li +4
Implementing large language models (LLMs)-driven root cause analysis (RCA) in cloud-native systems has become a key topic of modern software operations and maintenance. However, ex…
SpecPipe: Accelerating Pipeline Parallelism-based LLM Inference with Speculative Decoding
Haofei Yin, Mengbai Xiao, Tinghong Li +3
The demand for large language model inference is rapidly increasing. Pipeline parallelism offers a cost-effective deployment strategy for distributed inference but suffers from hig…