7 papers
Cross-region Model Training with Communication-Computation Overlapping and Delay Compensation
Ying Zhu, Yang Xu, Hongli Xu +3
Training large language models (LLMs) requires massive computational resources, often necessitating the aggregation of geographically distributed data centers (\ie, cross-region tr…
A Novel Hat-Shaped Device-Cloud Collaborative Inference Framework for Large Language Models
Zuan Xie, Yang Xu, Hongli Xu +2
Recent advancements in large language models (LLMs) have catalyzed a substantial surge in demand for LLM services. While traditional cloud-based LLM services satisfy high-accuracy…
Lightweight and Post-Training Structured Pruning for On-Device Large Lanaguage Models
Zihuai Xu, Yang Xu, Hongli Xu +3
Considering the hardware-friendly characteristics and broad applicability, structured pruning has emerged as an efficient solution to reduce the resource demands of large language…
Efficient Deployment of Large Language Models on Resource-constrained Devices
Zhiwei Yao, Yang Xu, Hongli Xu +2
Deploying Large Language Models (LLMs) on resource-constrained (or weak) devices presents significant challenges due to limited resources and heterogeneous data distribution. To ad…
Adaptive Parameter-Efficient Federated Fine-Tuning on Heterogeneous Devices
Jun Liu, Yunming Liao, Hongli Xu +3
Federated fine-tuning (FedFT) has been proposed to fine-tune the pre-trained language models in a distributed manner. However, there are two critical challenges for efficient FedFT…
Collaborative Inference for Large Models with Task Offloading and Early Exiting
Zuan Xie, Yang Xu, Hongli Xu +2
In 5G smart cities, edge computing is employed to provide nearby computing services for end devices, and the large-scale models (e.g., GPT and LLaMA) can be deployed at the network…