most citedEfficient Deployment of Large Language Models on Resource-constrained Devices

3 citations · 4 across the 12 of their papers we have counts for

collaborators

14 papers

cs.DC2025

DySTop

Yizhou Shi, Qianpiao Ma, Yan Xu +4

Federated Learning (FL) has emerged as a potential distributed learning paradigm that enables model training on edge devices (i.e., workers) while preserving data privacy. However,…

cs.LG2025

Resource-Efficient Federated Fine-Tuning Large Language Models for Heterogeneous Data

Jun Liu, Yunming Liao, Hongli Xu +1

Fine-tuning large language models (LLMs) via federated learning, i.e., FedLLM, has been proposed to adapt LLMs for various downstream applications in a privacy-preserving way. To r…

cs.LG2025

A Novel Hat-Shaped Device-Cloud Collaborative Inference Framework for Large Language Models

Zuan Xie, Yang Xu, Hongli Xu +2

Recent advancements in large language models (LLMs) have catalyzed a substantial surge in demand for LLM services. While traditional cloud-based LLM services satisfy high-accuracy…

cs.LG2025

Efficient Federated Fine-Tuning of Large Language Models with Layer Dropout

Shilong Wang, Jianchun Liu, Hongli Xu +2

Fine-tuning plays a crucial role in enabling pre-trained LLMs to evolve from general language comprehension to task-specific expertise. To preserve user data privacy, federated fin…

cs.DC2025

Collaborative Speculative Inference for Efficient LLM Inference Serving

Luyao Gao, Jianchun Liu, Hongli Xu +3

Speculative inference is a promising paradigm employing small speculative models (SSMs) as drafters to generate draft tokens, which are subsequently verified in parallel by the tar…

cs.LG2025

Lightweight and Post-Training Structured Pruning for On-Device Large Lanaguage Models

Zihuai Xu, Yang Xu, Hongli Xu +3

Considering the hardware-friendly characteristics and broad applicability, structured pruning has emerged as an efficient solution to reduce the resource demands of large language…