7 papers
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices
Mingyu Sun, Xiao Zhang, Shen Qu +5
Providing lossless inference services of LLMs on edge devices remains challenging, especially given the extremely tight memory budgets. The existing offloading techniques inevitabl…
Distributed Bilevel Optimization with Dual Pruning for Resource-limited Clients
Mingyi Li, Xiao Zhang, Ruisheng Zheng +4
With the development of large-scale models, traditional distributed bilevel optimization algorithms cannot be applied directly in low-resource clients. The key reason lies in the e…
Data-Free Continual Learning of Server Models in Model-Heterogeneous Cloud-Device Collaboration
Xiao Zhang, Zengzhe Chen, Yuan Yuan +5
The rise of cloud-device collaborative computing has enabled intelligent services to be delivered across distributed edge devices while leveraging centralized cloud resources. In t…
TAMO: Fine-Grained Root Cause Analysis via Tool-Assisted LLM Agent with Multi-Modality Observation Data in Cloud-Native Systems
Xiao Zhang, Qi Wang, Mingyi Li +4
Implementing large language models (LLMs)-driven root cause analysis (RCA) in cloud-native systems has become a key topic of modern software operations and maintenance. However, ex…
PE-MA: Parameter-Efficient Co-Evolution of Multi-Agent Systems
Yingfan Deng, Anhao Zhou, Yuan Yuan +3
Multi-Agent Systems have recently emerged as a promising paradigm for collaborative reasoning and solving complex tasks. However, the design of collaborative learning algorithms in…
DistrAttention: An Efficient and Flexible Self-Attention Mechanism on Modern GPUs
Haolin Jin, Mengbai Xiao, Yuan Yuan +4
The Transformer architecture has revolutionized deep learning, delivering the state-of-the-art performance in areas such as natural language processing, computer vision, and time s…