4 papers
A Novel Hat-Shaped Device-Cloud Collaborative Inference Framework for Large Language Models
Zuan Xie, Yang Xu, Hongli Xu +2
Recent advancements in large language models (LLMs) have catalyzed a substantial surge in demand for LLM services. While traditional cloud-based LLM services satisfy high-accuracy…
Lightweight and Post-Training Structured Pruning for On-Device Large Lanaguage Models
Zihuai Xu, Yang Xu, Hongli Xu +3
Considering the hardware-friendly characteristics and broad applicability, structured pruning has emerged as an efficient solution to reduce the resource demands of large language…
Efficient Deployment of Large Language Models on Resource-constrained Devices
Zhiwei Yao, Yang Xu, Hongli Xu +2
Deploying Large Language Models (LLMs) on resource-constrained (or weak) devices presents significant challenges due to limited resources and heterogeneous data distribution. To ad…
Collaborative Inference for Large Models with Task Offloading and Early Exiting
Zuan Xie, Yang Xu, Hongli Xu +2
In 5G smart cities, edge computing is employed to provide nearby computing services for end devices, and the large-scale models (e.g., GPT and LLaMA) can be deployed at the network…