3 papers
cs.LG2025
Lightweight and Post-Training Structured Pruning for On-Device Large Lanaguage Models
Zihuai Xu, Yang Xu, Hongli Xu +3
Considering the hardware-friendly characteristics and broad applicability, structured pruning has emerged as an efficient solution to reduce the resource demands of large language…
cs.DC2024
Adaptive Parameter-Efficient Federated Fine-Tuning on Heterogeneous Devices
Jun Liu, Yunming Liao, Hongli Xu +3
Federated fine-tuning (FedFT) has been proposed to fine-tune the pre-trained language models in a distributed manner. However, there are two critical challenges for efficient FedFT…
cs.DC2024
Collaborative Inference for Large Models with Task Offloading and Early Exiting
Zuan Xie, Yang Xu, Hongli Xu +2
In 5G smart cities, edge computing is employed to provide nearby computing services for end devices, and the large-scale models (e.g., GPT and LLaMA) can be deployed at the network…